Runtime detection, Slack alerts within seconds, native root cause, and one-click fixes. For every agent framework, in every language.
Not one failure mode — a category. Agents report success they didn't earn, get quietly downgraded to a weaker model, or hand work back and forth forever. Every run returns 200. Your logs look clean. Your customers don't.
Observability tools tell you what happened after your users complained. Dunetrace catches structural and semantic failures at runtime, stops the worst ones with policies, and hands you the root cause and a fix in Slack.
Not just an alerting bolt-on. Dunetrace covers everything from raw event capture to stopping a bad run before it finishes.
Self-hosted by default. Your prompts and tool arguments never leave your infrastructure. Structural detection runs entirely against your own Postgres — no LLM calls, no third-party services. The optional semantic evaluation worker is off by default; enabling it sends sampled run content to the LLM provider you configure.
Every run, tool call, and LLM exchange, the raw data everything else is built on.
34 zero-LLM detectors, 31 of them in-path, sub-500μs overhead†. Always on, no configuration.
LLM-based judgment: hallucination, task completion, cross-turn frustration. Post-hoc, sampling-based, opt-in.
Learn more →Policies that stop, redirect, or downgrade a run while it's happening.
Learn more →Native root-cause analysis, auto-applied policy fixes, or a one-click draft PR. No third-party tracer required.
† A per-event guideline, not a per-run guarantee. On the default HTTP ingest path the SDK adds roughly 10–25μs per event, so a ~20-event run stays under 500μs while larger runs scale proportionally. Policies that trigger on a signal run the detector suite in-path — measured ~130μs per event, roughly 10× the baseline. Architecture →
Detection tells you what broke; prevention stops the break from happening. Dunetrace's policy engine intervenes mid-run, blocking destructive tool calls, pausing high-value actions for human approval, switching failing models, and capping runaway iteration.
Destructive tool calls with fabricated arguments never execute.
High-value actions pause mid-run for human sign-off in Slack.
Runaway iteration stops before it drains your budget.
Most observability tools tell you what broke after your users complained. Dunetrace prevents the break at runtime.
Six categories of failure across structural and semantic layers. Every category can trigger a runtime policy that stops the failure mid-run.
A failure is detected and the Slack notification fires within seconds. It carries the root cause, the evidence, and one-click actions, before you ever open the dashboard.
Root cause, evidence, and suggested fix in the message itself. No dashboard round-trip needed to triage.
Apply fix, snooze pattern, mark false positive, or approve high-value operations without leaving Slack.
Per-detector severity thresholds, per-agent channel routing, and per-org integration configuration.
Dunetrace works alongside your existing evaluation tools. Pull your Langfuse, LangSmith, or Braintrust evaluation results into Dunetrace's issue view so everything lands in one dashboard. Or use Dunetrace's native semantic evaluation instead. Your choice.
Ask questions about your agents in natural language, from inside your editor. Dunetrace's MCP server loads full trace context, root cause, and fix suggestions directly into your coding agent.
pip install dunetrace-mcp.cursor/mcp.json (see docs).book_calendar() returned a 503, and the agent's next step told the user the meeting was confirmed.Instrument your agent in 5 minutes. Get your first alert before your next user does.