Documentation

Connect your monitoring

Two provider-agnostic methods, both available today. Either can be implemented in under an hour, and either is enough to move a requirement from Attested to Verified.

If you already run an OpenTelemetry collector, that is the route we expect most customers to take — see the OpenTelemetry guide (coming soon; you can configure it today, but the receiver is not accepting traffic yet). The webhook and heartbeat below work now, regardless of tooling. Start at the documentation hub or jump to troubleshooting.

What we read

We read monitoring metadata only: whether monitoring exists, when it last ran, error counts, and model or version changes. We never receive or store prompts, completions, customer data, or your agents' inputs and outputs. Everything else in the stream is discarded on receipt.

Ten signal types carry all of it. Each one is recorded as the fact it actually observed: a scheduler reporting in is not statistical drift, and we will not describe it as such to an underwriter.

Signal typeWhat it provesPromotes to Verified
drift_monitoringOngoing monitoring of model quality; Performance degradation detectionYes
monitoring_livenessA monitoring process exists and runs on a schedule; Configuration fingerprint is being watchedYes
error_rateError-rate obligations; Incident reporting and notificationYes
version_changeMaterial change notification; Model or version disclosureYes
configuration_changeMaterial change notification; Change control over deployed configurationYes
harness_scorePerformance and accuracy thresholds; Target Model Metric conditions on a quoteYes
human_review_rateReal-time human review of automated decisions; Qualified human oversight of a regulated serviceYes
guardrail_integrityA condition precedent that guardrails remain in force; Guardrails unchanged since inceptionYes
drift_cure_eventA cure obligation with a stated deadline; Drift discovered and closedYes
self_checkEvidence that a process exists — never evidence of an outcomeNo — documented only

self_check is evidence that a process exists, not that an outcome held, so it is filed and shown but never lifts a requirement to Verified.

Method 1 — Generic webhook (you push)

Create a webhook connector in the app. We issue a unique endpoint URL and a signing secret, shown once. POST the documented payload whenever your monitoring runs.

Payload

{
  "agent_ref": "patient-intake-prod",
  "signals": [
    {
      "type": "drift_monitoring",
      "observed_at": "2026-08-25T09:00:00Z",
      "summary": "Drift evaluation suite ran; 412 checks in the last 7 days",
      "payload": { "checks_7d": 412, "drift_detected": false }
    },
    {
      "type": "error_rate",
      "observed_at": "2026-08-25T09:00:00Z",
      "summary": "0.34% error rate over 18,422 runs; no incidents",
      "payload": { "runs": 18422, "errors": 63, "incidents": 0 }
    },
    {
      "type": "version_change",
      "observed_at": "2026-08-20T14:10:00Z",
      "summary": "Model changed from claude-3-5-sonnet to claude-sonnet-4",
      "payload": { "from": "claude-3-5-sonnet", "to": "claude-sonnet-4" }
    }
  ]
}
  • agent_ref must match the reference you mapped to an agent on the connector, otherwise we answer 409 rather than guessing. It may be omitted only when every signal in the body is an operational event (see the visibility connector below); a monitoring signal without it is refused with 400.
  • observed_at is optional and defaults to receipt time; send the real observation time where you have it.
  • error_rate and human_review_rate are refused as evidence when payload.window_complete is false: a window you could not read end to end is not a measurement. We answer 200, list the signal under signals_not_evidenced, and record nothing as evidence.
  • human_review_rate requires decisions and reviewed (integers, reviewed ≤ decisions). We store the share and both counts, so a carrier can re-derive the rate. A reported zero is a fact about your decisioning, not a missing reading.
  • guardrail_integrity requires intact (boolean) and guardrail_fingerprint. This closes a condition precedent, so it is watched across the term rather than signed once at inception.
  • drift_cure_event requires detected_at and cured_at (ISO 8601). We store the elapsed days against your cure period. Drift closed without an event is not visible to us and is not counted.
  • harness_score may carry payload.dimension_scores. Dimension names are matched case-insensitively, and a contract dimension we cannot find fails closed — it is never defaulted to a passing score. Without a validity window a harness score is treated as valid for 30 days, the ceiling the rail enforces.
  • summary is capped at 400 characters, up to 20 signals per request, 64KB per body, 120 deliveries per connector per hour.

Signature

Send x-riskborne-timestamp (Unix seconds) and sign <timestamp>.<raw body> with HMAC-SHA256, hex encoded, in x-riskborne-signature. Unsigned or mis-signed requests are rejected with 401 and recorded in your delivery log. A timestamp more than five minutes from our clock is refused with 401, and an exact re-send of a delivery we have already accepted is refused with 409 — the signed body is recorded per connector, so a captured request cannot be replayed.

Contract note, 2 September 2026. Body-only signing (no timestamp header) had no replay defense. It stays accepted until 9 September 2026 so nothing already wired breaks; those calls come back with a deprecation string in the 200 response. After that date they are refused with 401.

import { createHmac } from "node:crypto";

// Sign the RAW body bytes, exactly as sent. Do not re-serialise the object.
const timestamp = Math.floor(Date.now() / 1000).toString();
const signature = createHmac("sha256", process.env.RISKBORNE_SECRET)
  .update(`${timestamp}.${rawBody}`)
  .digest("hex");

await fetch(url, {
  method: "POST",
  headers: {
    "content-type": "application/json",
    "x-riskborne-timestamp": timestamp,
    "x-riskborne-signature": signature,
  },
  body: rawBody,
});

Worked example

SECRET='the signing secret shown once when you created the connector'
URL='https://riskborne.com/api/public/connector-webhook/<your-token>'
BODY='{"agent_ref":"patient-intake-prod","signals":[{"type":"drift_monitoring","summary":"Drift suite ran; 412 checks in the last 7 days","payload":{"checks_7d":412,"drift_detected":false}}]}'
TS=$(date +%s)

SIG=$(printf '%s.%s' "$TS" "$BODY" | openssl dgst -sha256 -hmac "$SECRET" -hex | sed 's/^.* //')

curl -X POST "$URL" \
  -H 'content-type: application/json' \
  -H "x-riskborne-timestamp: $TS" \
  -H "x-riskborne-signature: $SIG" \
  -d "$BODY"

# 200 {"ok":true,"signals_recorded":1}
# 409 {"ok":false,"error":"This exact signed delivery has already been accepted."}

Fixed test vector

If your signature does not reproduce this digest, the fault is in your signing code rather than your secret. Every character below is part of the signed string, including the single dot between the timestamp and the body.

secret     rb_test_secret_do_not_use
timestamp  1757340000
body       {"agent_ref":"patient-intake-prod","signals":[{"type":"monitoring_liveness","summary":"Monitoring ran"}]}

signed string   1757340000.{"agent_ref":"patient-intake-prod","signals":[{"type":"monitoring_liveness","summary":"Monitoring ran"}]}
HMAC-SHA256 hex 30e06ba883655cb51a261c033e02512adeedcc130b3d6c2c974f5c699c2395ba

A rejection now names which recipe your signature actually corresponds to — body without the timestamp prefix, the two joined without the dot, the timestamp last, or a re-serialised copy of your JSON. If none of them match, the secret itself is wrong and the message says so.

How to tell it is working

The connector page shows a “waiting for first signal” state that resolves the moment a delivery lands, and then displays the payload we received so you can confirm it matches what you sent. Requirements evidenced by that signal move to Verified in your next report view within seconds.

Harness scores — the full contract

A test harness reports a number against a threshold a carrier wrote into a quote. Four things are needed to emit one, and all four are below.

1 · Receiver contract

POST /api/public/connector-webhook/<connector-token>, content-type: application/json. One body carries agent_ref and between 1 and 20 signals, 64KB maximum, 120 deliveries per connector per hour.

  • Required for harness_score: type, a numeric score, and a non-empty harness naming the suite that produced it. A score without its suite is refused — an unnamed harness cannot be read against a contract.
  • Optional: observed_at (ISO 8601, defaults to receipt), summary (≤400 characters), valid_for_days (1–90, defaults to 30), and payload, where dimension_scores carries per-dimension numbers. A contract dimension the report omits fails closed; it is never defaulted to a pass.
  • Answers: 200 with signals_recorded, 401 for a bad signature or a timestamp more than five minutes from our clock, 409 for an unknown agent_ref or an exact replay.

2 · The key — two credentials, never one

The token in the URL says who is calling. The HMAC secret says who computed the bytes. They are separate values and must stay separate: a token appears in logs and proxies, a signing secret does not. The signing secret belongs to the connector, not to an agent: it is issued once when the connector is created and shared by every agent mapped to that connector, so rotating it changes the secret for all of them at once. Rotation is done from the connector page and keeps the endpoint URL unchanged; the secret is stored encrypted and shown only at issue and rotation. The output-data channel has its own secret again — a monitoring secret cannot post output rows and an output secret cannot post signals.

3 · Canonical signing rule

Covered bytes are the Unix-seconds timestamp, one ASCII full stop, then the raw request body exactly as transmitted — not a re-serialised copy, no whitespace changes, no key reordering. HMAC-SHA256, lowercase hex, sent in x-riskborne-signature with the same timestamp in x-riskborne-timestamp. Reproduce this vector offline before wiring anything; if it does not match, the fault is in the signing code.

The vector below is a static test artifact. Its timestamp is fixed so the digest is reproducible, which means it is far outside the five-minute window and will be answered 401 if you replay it against the live receiver. Use it to check your signing code offline, then sign a fresh timestamp for a real call — as the curl example does.

secret     rb_test_secret_do_not_use   (test secret — never a live one)
timestamp  1757538000   (fixed, offline only — see note)
raw body   275 bytes, exactly as shown, signed as sent

{"agent_ref":"riskborne-demo-h5-34-v1","signals":[{"type":"harness_score","harness":"riskborne-demo-h5-34-v1","score":0.264706,"observed_at":"2026-09-10T21:17:00Z","summary":"Hourly harness run","payload":{"threshold":0.7,"attribution":"Applicant-authored demo stand-in."}}]}

signed string   <timestamp> "." <raw body>
HMAC-SHA256 hex 59042299d3ec185738e6bbf4b49394822920e5b1d5d41457e9b7d0a7a628d12c
SECRET='the harness signing secret shown once for this agent'
URL='https://riskborne.com/api/public/connector-webhook/<connector-token>'
BODY='{"agent_ref":"riskborne-demo-h5-34-v1","signals":[{"type":"harness_score","harness":"riskborne-demo-h5-34-v1","score":0.264706,"observed_at":"2026-09-10T21:17:00Z","summary":"Hourly harness run","payload":{"threshold":0.7,"attribution":"Applicant-authored demo stand-in."}}]}'
TS=$(date +%s)

SIG=$(printf '%s.%s' "$TS" "$BODY" | openssl dgst -sha256 -hmac "$SECRET" -hex | sed 's/^.* //')

curl -X POST "$URL" \
  -H 'content-type: application/json' \
  -H "x-riskborne-timestamp: $TS" \
  -H "x-riskborne-signature: $SIG" \
  -d "$BODY"

# 200 {"ok":true,"signals_recorded":1}

4 · Connector identity

No new connector type is needed: a harness emits through the existing webhook connector. agent_ref must equal the reference mapped to the agent on that connector, and an unmapped reference is answered 409 rather than guessed at. The reference used in the worked example above, riskborne-demo-h5-34-v1, is example data only — no connector, agent or secret of that name has been issued. The connector page shows the mapped references and the token; the secret is shown once at creation and on rotation.

Three shapes the receiver will refuse: signal_type instead of type; a harness object instead of a string; and a threshold at the top of the signal. Put counts and the harness's own pass mark in payload — the threshold that decides a condition is the one written into the insurer's quote, never one sent with the score.

Evidence ceiling. A harness you run produces a documented result. Verified requires your insurer to name the suite in your contract. A score below its threshold is shown as a measured failure, not as a pending reading.

Method 2 — Agent heartbeat (we poll)

Expose one read-only HTTPS endpoint returning the document below. We poll it hourly with the bearer token you supply, and record every poll — successful or not.

{
  "agents": [
    {
      "id": "patient-intake-prod",
      "name": "Patient intake assistant",
      "model": "claude-sonnet-4",
      "version": "2026.08.3",
      "monitoring": {
        "enabled": true,
        "last_run": "2026-08-30T06:00:00Z",
        "type": "drift"
      },
      "errors_24h": 4,
      "errors_window_complete": true,
      "last_deployed": "2026-08-20T14:10:00Z"
    }
  ]
}
curl -H 'authorization: Bearer <the token you gave us>' \
     https://your-app.example.com/riskborne/heartbeat

Rules we apply

  • monitoring.enabled true with a last_run timestamp produces a drift_monitoring signal when type is drift or eval, and a monitoring_liveness signal when it is liveness — a scheduler reporting in is recorded as a scheduler reporting in.
  • errors_24h is no longer filed as evidence: a bare count has no denominator, so the error-rate obligation is closed by an error_rate signal pushed to the webhook. An absent or unconfirmed count is returned as a finding rather than a pass. config_changed_at produces a configuration_change signal, and last_deployed a version signal.
  • An endpoint that answers but reports no monitoring is a control gap, recorded as a finding — distinct from never having connected at all.
  • HTTPS only, no redirects followed, private, loopback, link-local and cloud metadata ranges rejected, DNS re-resolved at fetch time, short timeout, response size capped. Your endpoint’s raw body and errors are never surfaced to other users.

Model output data — a separate, consented channel

Everything above is metadata. The monitoring webhook and heartbeat never receive prompts, inputs or outputs, and that does not change. A carrier pricing an AI errors policy, however, needs the numbers themselves: what the model predicted, what turned out to be true, and what the difference cost. That travels on a different channel with its own consent record, its own credential and its own tables. Your monitoring secret cannot post to it, and its secret cannot post monitoring signals.

  • Endpoint /api/public/output-sample/<token>, issued per model from Model output data in the app, and only after consent is recorded.
  • Signed exactly like the monitoring webhook: HMAC-SHA256 over <timestamp>.<raw body>, lowercase hex, in x-riskborne-signature with x-riskborne-timestamp. Five-minute tolerance, replays rejected.
  • Six fields per record and no others: timestamp, optional customer reference, optional site reference, predicted value, ground truth value, optional loss. Each value is capped at 200 characters and a longer one is refused outright, never shortened — free text cannot fit through this channel.
  • At least 50 records per batch, from real deployment or an unseen representative set, never training data. Timestamps must be unique within a batch. A batch with any invalid row is rejected in full, reported row by row; nothing is partially stored.
  • Every batch declares where the true values came from. Self-reported values and values produced by a second model are capped at documented evidence and never presented as verified.
  • Withdraw consent and the endpoint stops accepting. Delete a dataset and every row of it goes with it.

The visibility connector — our own health, reported to you

RiskBorne’s own delivery infrastructure reports itself over the same signed webhook, on a separate connector with its own token and signing secret. That token is scoped to two event types and nothing else; a monitoring token cannot send them, and the visibility token cannot send monitoring signals. The scope boundary is the token’s allowlist, enforced at the receiver.

  • h5_emission_exception — our emitter attempted a delivery and the destination refused it. Optional fields: status_code (100–599), exception (≤400 characters), and response_body, capped at 1,000 characters at the receiver — a longer body is refused outright, never truncated silently.
  • emitter_liveness — a liveness reading of our emitter. state is delivering or degraded. Degraded means our emitter has not delivered in two consecutive opportunities. It is not a carrier determination, not a coverage outcome, and not a statement about your agent.
  • These events describe our operational state. They are stored as connection events, never as evidence signals; they are marked evidential = false in the render specification, so the promotion path refuses to read them even if one is ever mapped to a rule; and they render on the connection-health panel, never the evidence ladder. agent_ref is omitted for these events.
  • An optional top-level event_id (1–200 characters) makes an event idempotent: send the same event_id twice and the second delivery is accepted and ignored — a 200 with "duplicate": true, distinguishable from the 409 that answers an exact signature replay. Use a deterministic id (for example a hash of the event content) so a retry after a timeout cannot double-record.
  • A process cannot report its own silence, so a sender-side heartbeat is not the floor. Set an expected interval on the connector and the receiver holds it: if nothing arrives within twice that interval, the hourly sweep raises a finding on the connection-health panel itself.

Freshness — silence degrades

Every signal carries an expiry, capped by the requirement’s cure window. When a signal passes its expiry, the requirement it verified reverts to your attested position, the report says so plainly, and your workspace is notified. A pass never persists on the strength of an old signal.

Named providers

LangSmith, Langfuse and Datadog adapters are implemented against the documented provider APIs and are selectable, but no customer account has completed a live sync yet — test the connection before relying on one for a submission. AWS CloudWatch and Onyx are marked coming soon and cannot be selected. The webhook and heartbeat are the primary methods because they are the two that work for everyone.