“Agents now run for hours, call thousands of tools, and handle real money, real health data, and real customers,” said Zubin Koticha, chief executive of Raindrop. “When an agent fails, it does the wrong thing convincingly at scale until someone happens to notice.”
Raindrop has raised a Series A led by CRV. That takes total funding to $50m, the San Francisco company said on 17 September. It did not say how large the round itself is.
The company monitors AI agents once they are live. It reads production traces and flags what it calls semantic anomalies. Those are the hallucinated answers, tool misuse and behaviour changes that arrive with a model upgrade. Engineering teams see what changed, when it started and which users it hit.
Testing the change before it ships
Alongside the funding, Raindrop launched a product called Simulations, in research preview. It replays real production traffic and a team’s existing test cases against a proposed change to an agent. It then runs its anomaly detection over the results.
The pitch is a swipe at conventional evaluation. Traditional evals rely on test cases written in advance, Raindrop said. That means they mostly catch the failures a team already thought of. Simulations run on every pull request, and are meant to surface the behaviour changes nobody predicted.
Raindrop cited research from METR on how fast this is moving. The length of tasks agents can finish on their own roughly doubles every seven months. A single run can now last days and involve thousands of tool calls.
The company also claims Simulations gives companies outside the frontier labs the testing process those labs use internally. It points to OpenAI research on deployment simulation, which regenerates responses to de-identified production conversations with a candidate model. It points as well to Anthropic building synthetic universes to stress test its agents.
Who is backing it
Lightspeed Venture Partners and Y Combinator, both already on the cap table, joined the round. So did lead researchers from OpenAI, Anthropic and Thinking Machines, investing personally. The company also lists Figma Ventures and Vercel Ventures as backers.
Raindrop names Vercel, Framer and Clay as customers, alongside unnamed Fortune 100 enterprises in healthcare and logistics.
“Agents are fundamentally different from traditional software. They are highly capable, autonomous, and non-deterministic,” said Reid Christian, a general partner at CRV, in a statement. “Raindrop treats agent failure as a detection problem, the way a security company would.”
Bucky Moore, a partner at Lightspeed, said in a statement that bad behaviour will become catastrophic. He pointed to high-stakes settings such as defence.
“If we’re having an issue like a build failure or agents stuck in a loop, we see that issue in Slack,” said Bani Singh, an AI engineer at Vercel.
Koticha founded Raindrop with Ben Hylak and Alexis Gauba. The team includes engineers who built fraud detection models at Robinhood and anomaly detection at Square. Security engineers from Segment, Semgrep and Socket.dev have joined them.
A category getting crowded
Money keeps arriving in this corner. groundcover raised $100m in July to build observability for the AI era. Scaled Cognition took $100m from Khosla for reliable agents in June. Harvey bought Guardrails AI this month.
What separates the pitches is where they intervene. Raindrop is betting that agent reliability is a detection problem rather than a design one. It is also betting the useful moment is the pull request. Simulations sits in preview, so that bet is not settled yet.
Get the TNW newsletter
Get the most important tech news in your inbox each week.