HDG-443 — ReliAgent

Structured reliability middleware for AI agent tool calls — detects sycophantic gap-filling, hallucination-flavored language, tool-response contract violations, context degradation, repetition loops, confidence collapse, parameter drift, timeouts, and schema violations in real time.


Live Audience License Python


Table of Contents

  1. Description
  2. How It Works
  3. Detectors
  4. Usage
  5. Response Format
  6. Pricing
  7. Limitations
  8. Support & Contact
  9. License

Description

ReliAgent intercepts and validates a single AI agent tool-call record, detecting failure modes that plain text output hides from downstream pipelines.

Problem

AI agents fail silently. When an external tool returns a rate-limit error, an empty result set, a hallucinated value, or a malformed payload, the agent often receives all of these as undifferentiated text and has no built-in way to distinguish them. This produces sycophantic gap-filling (the agent invents plausible values), context degradation across multi-step workflows, hallucinated numeric claims inserted into structured pipelines, repetition loops that burn budget without making progress, and drifting or timing-out tool calls that go unnoticed until downstream logic breaks.

Solution

ReliAgent analyzes a tool call's parameters, response, and metadata, and returns a structured reliability report describing which failure modes were detected, with a confidence score and concrete remediation steps for each.

Target Audience


How It Works

  1. You send a single tool-call record via the /run endpoint: the tool name, the parameters that were passed to it, the raw response it returned, and the latency
  2. ReliAgent runs nine independent detectors against the record (see Detectors below) — some depend only on this one call, others also compare it against your recent call history
  3. Each detector returns a confidence score (0.0–1.0) and supporting evidence; results above the detection threshold are included in the report
  4. An overall status (pass, warn, fail, critical) is computed from the combined findings, along with deduplicated remediation recommendations

Important — what this does NOT do: - Does not call out to the tool itself — you provide the response that was already returned by your own tool call - Does not use any external LLM to judge your response — all detection is pattern/heuristic based (see Limitations) - Schema violation detection only runs if you provide an expected_schema in the request, and only checks a documented subset of JSON Schema — if your schema relies on a keyword outside that subset, you'll get a schema_partially_verified finding rather than a silent pass (see Limitations)

Call history and isolation: History-dependent detectors (context degradation, repetition loop, confidence collapse, parameter drift, and the adaptive timeout threshold) compare each call against your own recent call history, tracked per caller IP. This history is Redis-backed and persists across our service restarts — it does not reset on our end. Isolation between callers is IP-based: concurrent callers on distinct IPs never see each other's history, but a caller sharing an IP with others (e.g. behind the same NAT) shares a history pool with them. The adaptive timeout threshold has a 1,000ms floor — see the timeout row in Detectors for details.


Detectors

Detector Triggers on
sycophantic_gap_fill Weighted combination of: generic filler words ("certainly", "of course" — weak signal, cannot trigger alone), direct agreement/backtracking phrasing ("as you mentioned", "you're absolutely right"), unsupported consensus claims ("everyone knows", "it's widely accepted"), overclaiming/hedge-removal language ("this definitively proves", "without question"), more than 3 numeric values not traceable to your parameters, the response changing from your immediately preceding call to the same tool with identical parameters while agreement-flavored language is present, and/or your previous response to the same tool including citation markers (a URL, "source:", a numbered reference) that this response lacks despite similar unexplained numeric claims
hallucination Assertive knowledge-claim phrasing ("according to my training data", "I believe the answer is"...), or hedge language ("approximately", "roughly", "around", "about") co-occurring with 2 or more numeric values not traceable to your parameters
context_degradation Your reported context_window_used above 70% (warn) or 85% (high), and/or parameter keys present in your previous same-tool call but missing from this one
repetition_loop An identical response seen within your last 5 calls, and/or the same tool called 4+ times consecutively
confidence_collapse Your reported confidence_score below 0.5, a declining trend across your last 3 calls, a sharp rise (>0.2) from your immediately preceding call's score on a call that also trips sycophantic_gap_fill or repetition_loop, and/or uncertainty language in the response ("I'm not sure", "unclear", "cannot determine"...)
parameter_drift The same parameter keys present as your previous same-tool call, but a majority of the shared keys' values changed. A single changed value (e.g. an incrementing page number) does not trigger this — it requires both 2+ changed keys and a majority of shared keys changing
timeout Latency exceeding your supplied timeout_threshold_ms, or (if not supplied) an adaptive threshold of 3x your rolling average latency over the last 10 calls once you have that much history, or a 10-second default before then. Note: the adaptive threshold has a 1,000ms floor — it will never drop below 1,000ms regardless of your call history. If your tool typically responds in under ~300ms, your effective adaptive threshold is 1,000ms, not 3x your actual average. Use timeout_threshold_ms explicitly if you need a tighter threshold for fast tools
schema_violation Your response failing validation against your supplied expected_schema — see the supported keyword subset in Limitations
schema_partially_verified Your expected_schema relies (fully or partially) on a JSON Schema keyword this validator doesn't implement (anyOf, oneOf, allOf, not, format, $ref), and no violation was found in the portion that could be checked. This is a distinct finding from schema_violation — it means "we structurally can't verify this branch," not "we checked it and it's fine." Always scores warn, never fail/critical
response_integrity Structural contract/consistency defects in a single response: the response echoing back a different scalar value than a parameter you requested with, a success indicator co-occurring with a populated error field, a success indicator with an empty data/results/items/records payload, an error-shaped message/detail/description field with no explicit failure status, a count/total field that disagrees with the length of the list it describes, or an explicit truncated/has_more flag

A note on hallucination and response_integrity: ReliAgent only sees the tool call and its response — it never sees what the agent's next turn does with that response. That means it can't detect whether the agent's eventual answer invents facts, misinterprets valid data, or ignores the tool output entirely; those require observing the model's output after the tool call, which is outside this middleware's field of view. What it can detect, and what both of these detectors are built around, are the conditions in a tool response that are known to cause that kind of failure downstream — language in the response itself that reads as hallucination-flavored (hallucination), and structural defects in the response that are likely to force the next model turn to guess or gap-fill (response_integrity). Think of these as hallucination risk indicators grounded in this one call, not a claim about what the agent ultimately said.


Usage

curl -X POST https://app-443.hdgregory.com/run \
  -H "Content-Type: application/json" \
  -d '{
    "tool_name": "web_search",
    "parameters": {"query": "Q3 revenue figures for ACME Corp"},
    "response": {"results": [{"snippet": "Revenue was approximately $4.2 billion..."}]},
    "latency_ms": 312.4
  }'

Your first 15 calls are free — no signup or account needed, just solve a proof-of-work challenge first (see Pricing for the exact flow).

Request fields:

Field Type Required Description
tool_name string Yes Name of the tool that was called
parameters object Yes Parameters that were passed to the tool
response any Yes Raw response received from the tool
latency_ms float Yes Round-trip latency in milliseconds
call_id string No Your own identifier for this call; auto-generated if omitted
token_count integer No Token count of the response, if applicable
context_window_used float No Fraction of context window consumed (0.0–1.0)
confidence_score float No Your agent's own reported confidence in the response (0.0–1.0)
expected_schema object No A JSON schema to validate the response against (supported subset only — see Limitations)
timeout_threshold_ms float No Overrides the adaptive/default latency threshold used by the timeout detector for this call

Call history for repetition/context/drift/timeout detection is tracked automatically per caller IP — there's no field to set for this.


Response Format

{
  "success": true,
  "result": {
    "protocol_version": "1.1.0",
    "call_id": "a3f9c1d2",
    "tool_name": "web_search",
    "status": "warn",
    "timestamp": 1782090000.0,
    "failures": [
      {
        "failure_mode": "hallucination",
        "confidence": 0.65,
        "evidence": ["Hedge language co-occurs with 3 unexplained numeric value(s)"],
        "remediation": ["Request explicit source citations for all factual claims"]
      }
    ],
    "metrics": {
      "latency_ms": 312.4,
      "token_count": null,
      "context_window_used": null,
      "confidence_score": null,
      "failure_count": 1,
      "history_source": "redis"
    },
    "recommendations": ["Request explicit source citations for all factual claims"]
  },
  "error": null
}

status is one of pass, warn, fail, or critical. An empty failures array means the call passed all checks.

metrics.history_source reports how your call history was accessed for this call: "redis" under normal operation, or a "degraded_..." value in the rare case our Redis backend was unreachable — in which case history-dependent detectors ran without history for that single call rather than failing your request.


Pricing

Pay per validated call via x402 micropayments in USDC on Base.

Price $0.01 per call
Free trial Your first 15 calls are free — gated by a proof-of-work challenge, not by IP
Network Base mainnet, USDC

Claiming your free trial:

  1. GET /trial/challenge — returns a nonce and a required proof-of-work difficulty. No signup or payment needed for this step.
  2. Solve the challenge locally, at your own pace: find a solution string such that sha256(nonce + solution) has the required number of leading hex-zero characters (standard Hashcash-style proof-of-work). This typically takes anywhere from several seconds to under a minute on a single CPU core, depending on difficulty and how optimized your solver is.
  3. POST /trial/claim with {"nonce": ..., "solution": ...} — returns a trial_token good for your free-call budget (15 calls).
  4. Send your /run calls with an X-Trial-Token header set to that token. Once the budget is used, further calls automatically fall through to normal x402 payment — no separate step needed on your end.

Step 1 — get a challenge:

curl https://app-443.hdgregory.com/trial/challenge
# → {"nonce": "a3f9c1d2...", "difficulty": 6, "expires_in_s": 1800}

Step 2 — solve it locally (find a solution string where sha256(nonce + solution) starts with the returned number of hex zeros — read difficulty from Step 1's response rather than assuming it, since it's tunable server-side):

import hashlib, itertools, string

def solve(nonce, difficulty):
    chars = string.ascii_letters + string.digits
    for length in range(1, 20):
        for candidate in itertools.product(chars, repeat=length):
            solution = "".join(candidate)
            digest = hashlib.sha256((nonce + solution).encode()).hexdigest()
            if digest.startswith("0" * difficulty):
                return solution

nonce = "a3f9c1d2..."     # from Step 1
difficulty = 6            # from Step 1 — don't hardcode this, read it from the response
print(solve(nonce, difficulty))

Step 3 — claim your trial token:

curl -X POST https://app-443.hdgregory.com/trial/claim \
  -H "Content-Type: application/json" \
  -d '{"nonce": "a3f9c1d2...", "solution": "<your_solution>"}'
# → {"trial_token": "...", "calls_granted": 15, "expires_in_s": 86400}

Step 4 — use the token on /run calls:

curl -X POST https://app-443.hdgregory.com/run \
  -H "Content-Type: application/json" \
  -H "X-Trial-Token: <your_trial_token>" \
  -d '{"tool_name": "web_search", "parameters": {"query": "..."}, "response": {"...": "..."}, "latency_ms": 200.0}'

Each challenge is single-use, and each trial token has its own call budget and expiry — solving one challenge does not grant unlimited free calls.

No account or signup needed for either the free trial or paid use — payment is handled via the x402 protocol at the HTTP layer. Once your free calls are used, a request without a valid trial token or x402 payment returns an HTTP 402 Payment Required with the payment details needed to retry.


Limitations


Research

ReliAgent's detectors have been independently validated through Basanos, an open benchmarking study covering all eight detectors across 1,800 trials.

Report Description
Executive Summary 3-page overview of findings — verdict, results table, product defects surfaced and corrected
Basanos-2 Full Report Complete study — all eight detectors, 1,800 trials, methodology, confidence intervals
Basanos-1 Pilot Pilot phase — mechanism confirmation against six published benchmarks

Key findings: 0% false positive rate on seven of eight detectors; 100% true positive rate on all eight. Three genuine product defects were surfaced and corrected during the study. All data, harness code, and methodology are public at github.com/HDGForge-Labs/basanos.


Support & Contact

For questions or issues: [email protected]


License

Proprietary — © HD Gregory LLC. All rights reserved. See Terms of Service, Privacy Policy, and Refund Policy.