Docs / Architecture

Architecture

A core pipeline of services, one Postgres, one static dashboard, plus two optional Tier-2 workers. The SDK sends raw event data in-process, ingest writes events, the detector replays them, the explain layer templates a plain-English signal, and the alerts worker ships to Slack.

System overview

Dunetrace is a pipeline of independent services communicating through a shared Postgres database — no message broker. Each service does one job.

Agent Code
  └─▶ Dunetrace SDK       (raw event data → ingest events + OTel spans)
        └─▶ Ingest API    (POST /v1/ingest → Postgres, returns 202)
                ├─▶ Detector     (poll → RunState → 34 detectors → signals)
                ├─▶ Semantic     (optional, off by default → semantic signals)
                ├─▶ Integrations (optional, off by default → pulled evaluations)
                ├─▶ Alerts       (poll → explain → Slack / webhook)
                └─▶ Customer API (runs, signals, explanations → dashboard)
        ├─▶ stdout NDJSON (emit_as_json=True → Loki / Grafana Alloy)
        └─▶ OTel exporter (otel_exporter=… → Tempo / Honeycomb / Datadog)

Services

Ingest API · port 8001

The entry point for all SDK traffic. Its only job is to accept events as fast as possible and not lose them. It validates the schema, authenticates via the api_keys table, resolves the caller's org_id, writes the batch to Postgres, and only then returns 202 Accepted. A store failure or a row shortfall is a 503 with a Retry-After header instead, and the SDK keeps the batch and re-sends it whole. Nothing is ever acknowledged that was not written.

Why write before the 202? It used to be the other way round — 202 first, persistence in a background task, every failure swallowed into a log line — so a database outage looked like success to the SDK, which then discarded its only copy of the events. The extra round-trip is roughly one database write (~20ms) and it happens inside the SDK's background drain thread, never in your agent's own request path, so agent latency is unchanged. A re-sent batch is safe: already-stored event_ids are dropped on insert, so a batch whose 202 was lost in transit is accepted rather than duplicated.

POST /v1/otlp/traces is deliberately different: OTel exporters expect a fast 200 and retry on their own, so the OTLP receiver keeps its buffer-and-200 contract.

Detector worker

Background polling loop, every 5 seconds. The only process that runs detection logic.

  1. Fetches runs completed since last poll plus runs stalled longer than 90s
  2. Skips runs already in processed_runs
  3. Reconstructs RunState by replaying events
  4. Runs the Tier 1 detectors against the RunState. PROMPT_INJECTION_SIGNAL is handled by the SDK on raw input; the worker extracts the evidence from the run.started payload.
  5. Writes any FailureSignal rows
  6. UPSERTs the issues table for each fired signal and advances the clean-run counter; auto-resolves after 5 consecutive clean runs
  7. Marks the run processed

Why polling instead of streaming? A polling worker needs no message broker, survives restarts gracefully, and is trivial to reason about. At sub-100 runs/sec, 5-second latency is acceptable.

Poll watermark. Run discovery is bounded by a persisted per-shard watermark, so each poll scans only recent events rather than the entire retention window. This also keeps the monthly events partitions prunable — the partition key is received_at, and a query that never mentions it cannot prune. The window is expressed as "runs touched by a recent event" rather than "recent terminal events", which is what keeps late-arriving events triggering a re-detection. The watermark only advances after a poll that fully drains its backlog, so a busy or restarted worker can never step over unprocessed runs. Tune the re-scan overlap with WATERMARK_GRACE_SECS (default 3600).

Horizontal scaling. Set SHARD_COUNT=N and run N replicas with distinct SHARD_INDEX values; each polls only the runs whose agent_id hashes to its bucket. A misconfigured replica fails at startup rather than silently claiming no work.

Explain layer

A library — not a service. Imported by both the alerts worker and the customer API. Takes a FailureSignal, returns an Explanation in under 1 ms. Uses deterministic string templates, not LLM calls.

Three reasons for no LLM: latency (templates are instant), cost (zero per-signal API cost), consistency (same signal → same explanation, makes testing predictable).

Alerts worker

Background polling loop, every 10 seconds. The only process that sends external notifications. Fetches unalerted signals, computes rate context concurrently, calls explain(), formats for Slack Block Kit or webhook JSON, POSTs with exponential backoff up to 3 attempts. Marks alerted=TRUE only after at least one destination succeeds.

⚠
At-least-once delivery. If the worker crashes between sending and marking, the signal re-sends on restart. Receivers should treat (run_id, failure_type, detected_at) as the idempotency key.

Running more than one replica. Signals are claimed before delivery: the row is stamped in the same statement that selects it, using FOR UPDATE SKIP LOCKED. alerted can't serve that purpose on its own, because it is only set after a successful send — two workers scanning the same window would both see the signal outstanding and both deliver it. Claims expire after CLAIM_TIMEOUT_SECS (default 300) so a worker that dies mid-delivery doesn't strand its rows, and a poll that ends without delivering hands its claims back immediately.

To scale throughput, shard as well: SHARD_COUNT / SHARD_INDEX, or ALERTS_SHARD_COUNT / ALERTS_SHARD_INDEX to scale the alerts worker independently of the detector. Sharding on agent_id matters beyond throughput — one alert is sent per (org_id, agent_id, failure_type) group, so a group has to stay whole on a single worker.

Customer API · port 8002

FastAPI service — 100 routes across 25 routers. Powers the dashboard and any customer integrations. Signal responses include the full explanation inline.

It is not a read-only API. 42 of those routes are POST/PUT/PATCH/DELETE, and some of them change what a live agent does: POST /v1/policies can install a stop policy that terminates real runs as soon as the SDK next pulls policies, POST /v1/keys mints credentials, and POST /v1/approvals/{id}/decision releases a blocked agent. The endpoint table below is a selection, not the full surface.

Auth. Nearly every endpoint requires Authorization: Bearer <api_key> (skipped entirely in AUTH_MODE=dev, which is why dev mode is for a loopback-bound local stack only). Four routers are mounted without it, each for a stated reason: the pack catalog GET /v1/packs (static, no org context — the other pack routes authenticate individually), the Slack and Linear inbound webhooks (authenticated by the provider's own request signature rather than a Dunetrace key), and the GitHub App /callback (GitHub's own browser redirect, which carries no key — the other GitHub routes authenticate individually). The three probes below are unauthenticated by design and belong on an internal network only.

Writes are scope-gated on top of that. API keys carry scopes — ingest (the default an agent key gets), approve, and admin. admin gates every org-wide config write: API keys, policy writes and toggles, custom detectors, packs, org settings, and every integration config route. approve gates the approval decision, because the agent process being gated holds an ingest key and would otherwise be able to open its own gate. A key can mint at most the scopes it already holds, and an absent scope list fails closed to ingest-only. Reads stay ingest-accessible.

EndpointPurpose
GET /v1/agentsList agents with run counts, signal counts, failure breakdown
GET /v1/agents/{id}/runsPaginated run list — summary only
GET /v1/agents/{id}/signalsSignals with explanations; filters: severity, failure_type, include_shadow
GET /v1/agents/{id}/insightsAggregates — input patterns, daily trends, failure_rates, systemic_patterns
GET /v1/agents/{id}/issuesOpen/resolved issues per (agent, failure_type). Accepts optional status filter (open, resolved, reopened)
GET /v1/runs/{id}Full run — metadata, events, signals
POST /v1/signals/{id}/explainRoot-cause analysis, fully native — no request body, no external tracing system involved. Returns fix_category (dunetrace_native with a suggested_policy, or customer_code with fix_content/fix_patch), root_cause, apply_blocked. Requires one of ANTHROPIC_API_KEY, OPENAI_API_KEY or MISTRAL_API_KEY; API_LLM_PROVIDER pins which one, and a pinned provider with no key is an error rather than a silent fall-through to another vendor
POST /v1/signals/{id}/open-prFor code_change fixes only: opens a draft GitHub PR. Per-org GitHub App first, else legacy GITHUB_TOKEN/GITHUB_REPO. Edits the real file when source mapping resolves it; otherwise a summary file. Blocked for PROMPT_INJECTION_SIGNAL (403)
POST /v1/signals/{id}/record-copyRecord a clipboard-path fix in the fixes table
GET /v1/signals/{id}/fix-statusReturn fix history and recurrence verdict (verified / likely_fixed / still_occurring / insufficient_data)
GET /v1/agents/{id}/performance-trendsDaily structural/semantic signal rate, cost, latency over a 7/30/90-day window, plus failure-mode deltas and a self-baseline comparison
POST/GET/DELETE /v1/orgs/integrations/{github,langfuse,langsmith,braintrust,slack,linear}Connect, check, or remove a per-org integration
POST /v1/semantic-signalsPush an evaluation result from any external source, correlated via trace_id
POST/GET/DELETE /v1/agents/{id}/source-configExplicit repo/file mapping for GitHub PR source resolution
GET/POST/PATCH/DELETE /v1/custom-detectorsPlain-English custom detectors — preview, create (shadow mode), activate/pause, delete
GET /healthLiveness only — {"status":"ok","version":…}, no database round-trip; never fails while the process serves HTTP
GET /readyReadiness — 200 when the database answers at the schema version this build needs, else 503 with the same JSON body (db, schema_version, required, pool). What the compose healthchecks target
GET /metricsPrometheus exposition — request and LLM-call counts and latency, build and schema version. Unauthenticated, internal network only; every service exposes one, the workers on their own METRICS_PORT

See Semantic Evaluation and Integrations for the full endpoint reference on each of those surfaces.

Dashboard · port 3000

A single-page HTML app served by nginx. No build step — plain HTML/CSS/JS fetching from the Customer API. Auto-refreshes every 15 seconds. All data is computed client-side.

SDK framework integrations

Two first-class SDKs send events to the same ingest API — runs from either appear together in the dashboard under the same agent_id.

SDKInstallEntry point
Python (dunetrace)pip install dunetracefrom dunetrace import Dunetrace
TypeScript / Node.js (dunetrace)npm install dunetraceimport { Dunetrace } from "dunetrace"

The Python SDK ships framework integrations for LangChain, CrewAI, AutoGen, and Langfuse. The TypeScript SDK supports HTTP ingest and Loki NDJSON; OTel spans are Python-only.

FrameworkClassInstall
LangChain / LangGraphDunetraceCallbackHandlerpip install 'dunetrace[langchain]'
CrewAI 1.xDunetraceCrewCallbackpip install dunetrace crewai
AutoGen (autogen-agentchat ≥ 0.4)DunetraceAutoGenObserverpip install dunetrace autogen-agentchat autogen-ext
OpenLLMetry / OTel receiverDunetraceOTelReceiverpip install 'dunetrace[otel]'

SDK output modes

Three independent output paths that can be combined:

ModeHow to enableDestination
HTTP ingest (default)endpoint="http://…"Ingest API → Postgres → Detector
Loki NDJSONemit_as_json=Truestdout → Promtail/Alloy → Loki
OTel spansotel_exporter=DunetraceOTelExporter(provider)OTel collector → Tempo / Honeycomb / Datadog

All three can be active at once. OTel and NDJSON are zero-cost when disabled. Pass endpoint=None for OTel-only or Loki-only deployments.

Database schema

Seven tables. Event payloads land in the payload JSONB column; when you self-host, that data never leaves your own PostgreSQL.

CREATE TABLE events (
    id             BIGSERIAL PRIMARY KEY,
    batch_id       TEXT             NOT NULL,
    event_type     TEXT             NOT NULL,
    run_id         TEXT             NOT NULL,
    agent_id       TEXT             NOT NULL,
    agent_version  TEXT             NOT NULL,
    step_index     INTEGER          NOT NULL,
    timestamp      DOUBLE PRECISION NOT NULL,
    payload        JSONB            NOT NULL,
    parent_run_id  TEXT,
    received_at    TIMESTAMPTZ      NOT NULL DEFAULT NOW()
);

CREATE TABLE failure_signals (
    id             BIGSERIAL PRIMARY KEY,
    failure_type   TEXT        NOT NULL,
    severity       TEXT        NOT NULL,
    run_id         TEXT        NOT NULL,
    agent_id       TEXT        NOT NULL,
    agent_version  TEXT        NOT NULL,
    step_index     INTEGER     NOT NULL,
    confidence     REAL        NOT NULL,
    evidence       JSONB       NOT NULL,
    detected_at    TIMESTAMPTZ NOT NULL DEFAULT NOW(),
    shadow         BOOLEAN     NOT NULL DEFAULT TRUE,
    alerted        BOOLEAN     NOT NULL DEFAULT FALSE
);

CREATE TABLE processed_runs (…);
CREATE TABLE api_keys (…);
CREATE TABLE issues (…);
CREATE TABLE digest_log (…);

CREATE TABLE fixes (
    id                    BIGSERIAL    PRIMARY KEY,
    run_id                TEXT         NOT NULL,
    signal_id             BIGINT       NOT NULL,
    fix_content           TEXT         NOT NULL,
    fix_type              TEXT         NOT NULL DEFAULT 'prompt_addition',
    applied_via           TEXT         NOT NULL,   -- 'github_pr' or 'clipboard'
    langfuse_prompt_name  TEXT,                    -- historical column, from the removed Langfuse apply-fix flow — always NULL now
    langfuse_version      INTEGER,                 -- historical column, repurposed to store the GitHub PR number when applied_via='github_pr'
    applied_at            TIMESTAMPTZ  NOT NULL DEFAULT NOW()
);

Performance

ComponentLatencyThroughput
SDK _emit()<1 μsMillions/sec
SDK drain thread200 ms idle poll100 events/batch
Ingest API (202)~5 ms~1,000 req/sec
Detector poll cycle5 s~100 runs/cycle
Explain layer<1 mssynchronous
Alerts poll cycle10 s50 signals/cycle
Customer API~10 ms~500 req/sec

Agent overhead: under 500 μs per run with default HTTP ingest. The drain thread is entirely background. Even under backpressure (ingest API down), the ring buffer drops the oldest events rather than blocking the agent.

Failure modes

  • Ingest API down — drain thread drops unshippable events. Agent never blocks. Events during outage are lost.
  • Detector worker down — runs queue up. When the worker restarts, it catches up. Signals delayed but not lost.
  • Postgres down — ingest returns 503. SDK buffers up to 10,000 events, then rolls. Observability data loss is acceptable during DB outages.
  • Alerts worker down — signals accumulate as alerted=FALSE. On restart, delivery resumes. At-least-once.