Gremlin Foresight AI finds reliability risks, tests fixes
Gremlin has made Gremlin Foresight AI generally available following a beta phase. The tool searches production systems for reliability risks, suggests a fix, and then reruns the original test to confirm the fix did its job.
The company is targeting engineering teams that now ship code faster with AI assistance and worry that the extra speed will bring more outages, degraded services and unhappy customers.
Faster code, more chances to fail
Gremlin CEO Kolton Andrus sees the risk growing alongside the speed.
"AI-driven development means shipping code at 10X velocity; it also means 10x the opportunity for bugs, risks, and failures," he said.
Andrus also compared the product to the AI SRE tools now on the market. SRE stands for site reliability engineering, the practice of keeping online services running. He said these tools help teams respond to incidents more quickly, but they still kick in only after something has failed.
"While the rise of AI SRE tools are great for helping teams respond to incidents faster, it's still cleanup after something breaks. Gremlin Foresight AI finds risks proactively, delivers the fix, and verifies the fix worked by rerunning the test. This provides teams with the assurances needed to move confidently," Andrus explained.
Built on a decade of failure data
The tool is built on Gremlin's proprietary Failure Atlas, a collection of more than ten years of cause-and-effect data on how online systems break. According to Gremlin, this means its recommendations come from real failure patterns and operational experience rather than generic best practices.
Each fix is checked against the same test that first exposed the problem, so the process runs from spotting a weakness all the way to showing it is gone.
Gremlin lists four main capabilities:
- Proactive risk detection: finding weaknesses and failure conditions before they turn into incidents
- Guided remediation: recommending and delivering a specific fix, either as a configuration patch or an infrastructure-as-code change
- Continuous validation: rerunning the original test to confirm the fix works, and repeating tests as systems change
- Measurable resilience: tracking reliability scores across services and teams to measure progress and decide where to invest next
From Chaos Monkey to agentic resilience
Andrus has worked on this problem for a long time. Before starting Gremlin, he was responsible for the uptime of Amazon's retail site. He later moved to Netflix, where he built the company's second generation of fault-injection tooling after Netflix's open source project Chaos Monkey became popular. Fault injection means deliberately breaking parts of a system to see how the rest of it copes.
Gremlin says it has since taken Chaos Engineering beyond one-off experiments and turned it into a broader approach to reliability management. Its platform lets teams run planned experiments, limit the "blast radius" of each test, and use continuous automated testing to check that fixes keep working over time. Reliability scores give the whole organization a shared standard and give leaders visibility over where to direct investment and how to hold teams accountable.
Mike Dauber, General Partner at Amplify Partners, said many teams lacked the time or expertise to run such experiments consistently.
"Foresight AI is like a trainer who does the reps for you. Your systems get stronger without your team doing all the manual work," Dauber said.
Gremlin describes the result as "agentic resilience": it spots reliability risks earlier and automates fixes where appropriate.
Our Take
Foresight AI fits a pattern we keep seeing. Vendors are no longer satisfied with AI that only finds problems. They want it to fix them as well, much like Legit Security's agent for open-source dependencies. Gremlin's choice to rerun the original test is worth noting, because an unverified automated fix can create new failures of its own.
For security teams, availability is part of the job, and stress-testing systems before attackers or bad deployments do overlaps with controlled AI penetration testing. It is worth watching how much control organizations are willing to give an agent that pushes infrastructure changes to production, and whether they will want ways to rewind AI agent actions when something goes wrong. Gremlin's claims about the Failure Atlas will also need independent results from customers before they can be judged.
