Structured reliability middleware for AI agent tool calls — detects sycophantic gap-filling, hallucination-flavored language, tool-response contract violations, context degradation, repetition loops, confidence collapse, parameter drift, timeouts, and schema violations in real time.
ReliAgent intercepts and validates a single AI agent tool-call record, detecting failure modes that plain text output hides from downstream pipelines.
AI agents fail silently. When an external tool returns a rate-limit error, an empty result set, a hallucinated value, or a malformed payload, the agent often receives all of these as undifferentiated text and has no built-in way to distinguish them. This produces sycophantic gap-filling (the agent invents plausible values), context degradation across multi-step workflows, hallucinated numeric claims inserted into structured pipelines, repetition loops that burn budget without making progress, and drifting or timing-out tool calls that go unnoticed until downstream logic breaks.
ReliAgent analyzes a tool call's parameters, response, and metadata, and returns a structured reliability report describing which failure modes were detected, with a confidence score and concrete remediation steps for each.
/run endpoint: the tool name, the parameters that were passed to it, the raw response it returned, and the latencypass, warn, fail, critical) is computed from the combined findings, along with deduplicated remediation recommendationsImportant — what this does NOT do:
- Does not call out to the tool itself — you provide the response that was already returned by your own tool call
- Does not use any external LLM to judge your response — all detection is pattern/heuristic based (see Limitations)
- Schema violation detection only runs if you provide an expected_schema in the request, and only checks a documented subset of JSON Schema — if your schema relies on a keyword outside that subset, you'll get a schema_partially_verified finding rather than a silent pass (see Limitations)
Call history and isolation: History-dependent detectors (context degradation, repetition loop, confidence collapse, parameter drift, and the adaptive timeout threshold) compare each call against your own recent call history, tracked per caller IP. This history is Redis-backed and persists across our service restarts — it does not reset on our end. Isolation between callers is IP-based: concurrent callers on distinct IPs never see each other's history, but a caller sharing an IP with others (e.g. behind the same NAT) shares a history pool with them. The adaptive timeout threshold has a 1,000ms floor — see the timeout row in Detectors for details.
| Detector | Triggers on |
|---|---|
sycophantic_gap_fill |
Weighted combination of: generic filler words ("certainly", "of course" — weak signal, cannot trigger alone), direct agreement/backtracking phrasing ("as you mentioned", "you're absolutely right"), unsupported consensus claims ("everyone knows", "it's widely accepted"), overclaiming/hedge-removal language ("this definitively proves", "without question"), more than 3 numeric values not traceable to your parameters, the response changing from your immediately preceding call to the same tool with identical parameters while agreement-flavored language is present, and/or your previous response to the same tool including citation markers (a URL, "source:", a numbered reference) that this response lacks despite similar unexplained numeric claims |
hallucination |
Assertive knowledge-claim phrasing ("according to my training data", "I believe the answer is"...), or hedge language ("approximately", "roughly", "around", "about") co-occurring with 2 or more numeric values not traceable to your parameters |
context_degradation |
Your reported context_window_used above 70% (warn) or 85% (high), and/or parameter keys present in your previous same-tool call but missing from this one |
repetition_loop |
An identical response seen within your last 5 calls, and/or the same tool called 4+ times consecutively |
confidence_collapse |
Your reported confidence_score below 0.5, a declining trend across your last 3 calls, a sharp rise (>0.2) from your immediately preceding call's score on a call that also trips sycophantic_gap_fill or repetition_loop, and/or uncertainty language in the response ("I'm not sure", "unclear", "cannot determine"...) |
parameter_drift |
The same parameter keys present as your previous same-tool call, but a majority of the shared keys' values changed. A single changed value (e.g. an incrementing page number) does not trigger this — it requires both 2+ changed keys and a majority of shared keys changing |
timeout |
Latency exceeding your supplied timeout_threshold_ms, or (if not supplied) an adaptive threshold of 3x your rolling average latency over the last 10 calls once you have that much history, or a 10-second default before then. Note: the adaptive threshold has a 1,000ms floor — it will never drop below 1,000ms regardless of your call history. If your tool typically responds in under ~300ms, your effective adaptive threshold is 1,000ms, not 3x your actual average. Use timeout_threshold_ms explicitly if you need a tighter threshold for fast tools |
schema_violation |
Your response failing validation against your supplied expected_schema — see the supported keyword subset in Limitations |
schema_partially_verified |
Your expected_schema relies (fully or partially) on a JSON Schema keyword this validator doesn't implement (anyOf, oneOf, allOf, not, format, $ref), and no violation was found in the portion that could be checked. This is a distinct finding from schema_violation — it means "we structurally can't verify this branch," not "we checked it and it's fine." Always scores warn, never fail/critical |
response_integrity |
Structural contract/consistency defects in a single response: the response echoing back a different scalar value than a parameter you requested with, a success indicator co-occurring with a populated error field, a success indicator with an empty data/results/items/records payload, an error-shaped message/detail/description field with no explicit failure status, a count/total field that disagrees with the length of the list it describes, or an explicit truncated/has_more flag |
A note on hallucination and response_integrity: ReliAgent only sees the tool call and its response — it never sees what the agent's next turn does with that response. That means it can't detect whether the agent's eventual answer invents facts, misinterprets valid data, or ignores the tool output entirely; those require observing the model's output after the tool call, which is outside this middleware's field of view. What it can detect, and what both of these detectors are built around, are the conditions in a tool response that are known to cause that kind of failure downstream — language in the response itself that reads as hallucination-flavored (hallucination), and structural defects in the response that are likely to force the next model turn to guess or gap-fill (response_integrity). Think of these as hallucination risk indicators grounded in this one call, not a claim about what the agent ultimately said.
curl -X POST https://app-443.hdgregory.com/run \
-H "Content-Type: application/json" \
-d '{
"tool_name": "web_search",
"parameters": {"query": "Q3 revenue figures for ACME Corp"},
"response": {"results": [{"snippet": "Revenue was approximately $4.2 billion..."}]},
"latency_ms": 312.4
}'
Your first 15 calls are free — no signup or account needed, just solve a proof-of-work challenge first (see Pricing for the exact flow).
Request fields:
| Field | Type | Required | Description |
|---|---|---|---|
tool_name |
string | Yes | Name of the tool that was called |
parameters |
object | Yes | Parameters that were passed to the tool |
response |
any | Yes | Raw response received from the tool |
latency_ms |
float | Yes | Round-trip latency in milliseconds |
call_id |
string | No | Your own identifier for this call; auto-generated if omitted |
token_count |
integer | No | Token count of the response, if applicable |
context_window_used |
float | No | Fraction of context window consumed (0.0–1.0) |
confidence_score |
float | No | Your agent's own reported confidence in the response (0.0–1.0) |
expected_schema |
object | No | A JSON schema to validate the response against (supported subset only — see Limitations) |
timeout_threshold_ms |
float | No | Overrides the adaptive/default latency threshold used by the timeout detector for this call |
Call history for repetition/context/drift/timeout detection is tracked automatically per caller IP — there's no field to set for this.
{
"success": true,
"result": {
"protocol_version": "1.1.0",
"call_id": "a3f9c1d2",
"tool_name": "web_search",
"status": "warn",
"timestamp": 1782090000.0,
"failures": [
{
"failure_mode": "hallucination",
"confidence": 0.65,
"evidence": ["Hedge language co-occurs with 3 unexplained numeric value(s)"],
"remediation": ["Request explicit source citations for all factual claims"]
}
],
"metrics": {
"latency_ms": 312.4,
"token_count": null,
"context_window_used": null,
"confidence_score": null,
"failure_count": 1,
"history_source": "redis"
},
"recommendations": ["Request explicit source citations for all factual claims"]
},
"error": null
}
status is one of pass, warn, fail, or critical. An empty failures array means the call passed all checks.
metrics.history_source reports how your call history was accessed for this call: "redis" under normal operation, or a "degraded_..." value in the rare case our Redis backend was unreachable — in which case history-dependent detectors ran without history for that single call rather than failing your request.
Pay per validated call via x402 micropayments in USDC on Base.
| Price | $0.01 per call |
| Free trial | Your first 15 calls are free — gated by a proof-of-work challenge, not by IP |
| Network | Base mainnet, USDC |
Claiming your free trial:
GET /trial/challenge — returns a nonce and a required proof-of-work difficulty. No signup or payment needed for this step.solution string such that sha256(nonce + solution) has the required number of leading hex-zero characters (standard Hashcash-style proof-of-work). This typically takes anywhere from several seconds to under a minute on a single CPU core, depending on difficulty and how optimized your solver is.POST /trial/claim with {"nonce": ..., "solution": ...} — returns a trial_token good for your free-call budget (15 calls)./run calls with an X-Trial-Token header set to that token. Once the budget is used, further calls automatically fall through to normal x402 payment — no separate step needed on your end.Step 1 — get a challenge:
curl https://app-443.hdgregory.com/trial/challenge
# → {"nonce": "a3f9c1d2...", "difficulty": 6, "expires_in_s": 1800}
Step 2 — solve it locally (find a solution string where sha256(nonce + solution) starts with the returned number of hex zeros — read difficulty from Step 1's response rather than assuming it, since it's tunable server-side):
import hashlib, itertools, string
def solve(nonce, difficulty):
chars = string.ascii_letters + string.digits
for length in range(1, 20):
for candidate in itertools.product(chars, repeat=length):
solution = "".join(candidate)
digest = hashlib.sha256((nonce + solution).encode()).hexdigest()
if digest.startswith("0" * difficulty):
return solution
nonce = "a3f9c1d2..." # from Step 1
difficulty = 6 # from Step 1 — don't hardcode this, read it from the response
print(solve(nonce, difficulty))
Step 3 — claim your trial token:
curl -X POST https://app-443.hdgregory.com/trial/claim \
-H "Content-Type: application/json" \
-d '{"nonce": "a3f9c1d2...", "solution": "<your_solution>"}'
# → {"trial_token": "...", "calls_granted": 15, "expires_in_s": 86400}
Step 4 — use the token on /run calls:
curl -X POST https://app-443.hdgregory.com/run \
-H "Content-Type: application/json" \
-H "X-Trial-Token: <your_trial_token>" \
-d '{"tool_name": "web_search", "parameters": {"query": "..."}, "response": {"...": "..."}, "latency_ms": 200.0}'
Each challenge is single-use, and each trial token has its own call budget and expiry — solving one challenge does not grant unlimited free calls.
No account or signup needed for either the free trial or paid use — payment is handled via the x402 protocol at the HTTP layer. Once your free calls are used, a request without a valid trial token or x402 payment returns an HTTP 402 Payment Required with the payment details needed to retry.
sycophantic_gap_fill) detects that an answer changed with no change in input — it cannot verify which of the two responses, if either, was factually correctsycophantic_gap_fill) only looks for source-like markers — a URL, source:, citation:, or a [n] reference — disappearing between consecutive calls to the same tool. It does not verify that a present citation is real, relevant, or actually supports the claim it's attached toconfidence_collapse) requires both a sharp rise in your reported confidence_score from the previous call and corroborating evidence from sycophantic_gap_fill or repetition_loop on the same call — a rising confidence score alone is never flagged, since that's normal and expected in most sessionsresponse_integrity only runs its checks against dict-shaped responses; a bare string or number response yields no findings from this detector, since none of its checks (parameter-echo, status/error contradiction, count mismatch, etc.) apply to non-dict dataresponse_integrity does not check response freshness/staleness (cache age, TTL expiry) — that would require a timestamp field this schema doesn't currently carry. If this matters for your use case, let us know; it's a natural extensiontype, required, properties, enum, const, items (including recursive array-element validation), minimum, maximum, minLength, maxLength, minItems, maxItems, pattern, additionalProperties. Not supported: anyOf, oneOf, allOf, not, format, $ref. If your expected_schema relies (fully or partially) on one of these unsupported keywords and no violation is found under the supported portion, the response is not reported as a clean pass — you'll get a distinct schema_partially_verified finding instead (see Detectors), naming exactly which keyword(s) at which path couldn't be evaluated. That finding means "this branch was never actually checked," not "this branch is compliant"ReliAgent's detectors have been independently validated through Basanos, an open benchmarking study covering all eight detectors across 1,800 trials.
| Report | Description |
|---|---|
| Executive Summary | 3-page overview of findings — verdict, results table, product defects surfaced and corrected |
| Basanos-2 Full Report | Complete study — all eight detectors, 1,800 trials, methodology, confidence intervals |
| Basanos-1 Pilot | Pilot phase — mechanism confirmation against six published benchmarks |
Key findings: 0% false positive rate on seven of eight detectors; 100% true positive rate on all eight. Three genuine product defects were surfaced and corrected during the study. All data, harness code, and methodology are public at github.com/HDGForge-Labs/basanos.
For questions or issues: [email protected]
Proprietary — © HD Gregory LLC. All rights reserved. See Terms of Service, Privacy Policy, and Refund Policy.