66.30%
Multi-agent shared evidence / independent reasoning
13 → 4 paid web-search actions · 8 → 10 model Responses
$0.16499684 → $0.05560972 · avoided $0.10938712 gross
Eight department agents across two tenant identities performed useful current ecosystem research. SeenRelay shared one same-tenant evidence packet per organization while preserving all eight independent role analyses. Gross provider cost fell from $0.16499684 to $0.05560972. SeenRelay OpsCo was first-party real operating work; the second tenant was explicitly synthetic. Local coordination/integration overhead was not monetized, so this is not customer ROI, net customer savings, external adoption or a universal 66% claim.
First-party real operating workload + synthetic external-client rehearsal · baseline quality TOOL_NECESSITY_PASS_BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 1 complete paired run
Case study →
74.95%
Hosted compute / 64 GB tool session
4 → 1 provider executions
$1.9241253 → $0.4820025 · avoided $1.4421228
Mechanism-only benchmark: the exact task was deterministic SHA-256 and could be done locally. This row proves duplicate hosted-resource coalescing, not that a 64 GB container was necessary for SHA-256.
Controlled first-party benchmark · baseline quality MECHANISM_ONLY_LOCAL_WINS · semantic PASS · repeatability 3/3
Integration path →
66.88% mean
Agentic web search — explicit shared-snapshot contract
4 → 1 provider executions
4 top-level executions → 1 top-level execution · avoided 42.84%–85.73%
Mechanism-only benchmark for this exact task: all callers accepted one shared fresh provider-search snapshot, but the requested latest PyPI package version is available from PyPI directly. This proves shared-search coordination, not web-search necessity. Other agentic-search workloads may require k>1 or k=N.
Controlled first-party benchmark · baseline quality MECHANISM_ONLY_SOURCE_NATIVE_WINS · semantic PASS · repeatability 3/3
Integration path →
66.7%
Decision-time freshness
6 → 2 provider executions
6 provider reads → 2 provider reads · avoided 4 provider reads
Compared with Firecrawl provider-native maxAge on the same controlled schedule: native path stayed at 6 calls / 6 credits, while SeenRelay used 2 calls / 2 credits in 3/3 runs. Hard max-age policy remained satisfied; no claim of hidden source-change detection inside the allowed age window.
Controlled first-party benchmark · baseline quality BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
75%
Exact web extraction
4 → 1 provider executions
4 Firecrawl credits → 1 Firecrawl credit · avoided 3 Firecrawl credits
Exact same-boundary markdown extraction. A separate structured JSON contract failed semantic equality and remains a negative result.
Controlled first-party benchmark · baseline quality BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
75%
Weather downstream analysis
4 → 1 provider executions
4 paid analyses → 1 paid analysis · avoided 3 paid analyses
Mechanism-only benchmark: real Open-Meteo state stayed unchanged, but the paid classifier applied fixed rules that the harness itself could compute locally. This proves state-keyed recomputation control, not best-baseline weather economics.
Controlled first-party benchmark · baseline quality MECHANISM_ONLY_LOCAL_WINS · semantic PASS · repeatability 3/3
Integration path →
50%
Weather state transition
4 → 2 provider executions
4 paid analyses → 2 paid analyses · avoided 2 paid analyses
Controlled A,A,B,B transition forced a new analysis at B and recorded zero stale reuse after the change.
Controlled first-party benchmark · baseline quality BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
75%
Retail extraction guarded by price state
4 → 1 provider executions
4 Firecrawl credits → 1 Firecrawl credit · avoided 3 Firecrawl credits
Native-first conditional: the direct Apple read already resolves the monitored price/capacity state. Paid Firecrawl extraction is relevant only when the caller needs the richer extraction artifact; if state alone is sufficient, SeenRelay should self-reject.
Controlled first-party benchmark · baseline quality NATIVE_FIRST_CONDITIONAL · semantic PASS · repeatability 3/3
Integration path →
50%
Retail price / availability transition
4 → 2 provider executions
4 paid analyses → 2 paid analyses · avoided 2 paid analyses
Controlled authoritative A,A,B,B retail fixture. The changed price/availability fingerprint forced a new analysis before reuse resumed; this is not a live Apple price-change claim.
Controlled first-party benchmark · baseline quality BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
75%
Inventory / availability downstream extraction
4 → 1 provider executions
4 Firecrawl credits → 1 Firecrawl credit · avoided 3 Firecrawl credits
Real Adafruit inventory state. Firecrawl provider-native maxAge also stayed at 4 calls / 4 credits; SeenRelay used the retailer Product API as a cheap opaque state token and performed one paid extraction. If that native API already answers the entire need, SeenRelay should not be used.
Controlled first-party benchmark · baseline quality NATIVE_FIRST_CONDITIONAL · semantic PASS · repeatability 3/3
Integration path →
51.58% mean
Inventory state transition
4 → 2 provider executions
4 paid analyses → 2 paid analyses · avoided 2 paid analyses
Controlled A,A,B,B stock/availability transition. Each changed state forced a new paid analysis before reuse resumed; gross modeled reduction ranged from 50% to 53.1%. This is not a live retailer stock-change claim.
Controlled first-party benchmark · baseline quality BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
74.84%
AI gateway cold concurrent cache misses
4 → 1 provider executions
4 upstream executions → 1 upstream execution · avoided 3 upstream executions
Mechanism-only benchmark: against LiteLLM Proxy 1.103.2 + Redis exact cache, four cold concurrent misses caused four upstream executions and SeenRelay single-flight caused one. The exact benchmark output was a fixed constant, so this isolates the cold-concurrency gap rather than proving model-call necessity.
Controlled first-party benchmark · baseline quality MECHANISM_ONLY_LOCAL_WINS · semantic PASS · repeatability 3/3
Integration path →
79.61% mean
News / event downstream briefing
4 → 1 provider executions
4 paid briefings → 1 paid briefing · avoided $0.0546–$0.1010 modeled
Real Hacker News event authority plus real OpenAI web-search briefing. In 3/3 stable-state runs, four paid briefings became one and gross avoided cost was $0.05462206–$0.10098029 per group (70.95%–90.00%). HN remained the native state authority; independent customer prevalence is not claimed.
Controlled first-party benchmark · baseline quality TOOL_NECESSITY_PASS_BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
49.93% mean
News / event state transition
4 → 2 provider executions
4 paid briefings → 2 paid briefings · avoided $0.0116–$0.0664 modeled
Controlled A,A,B,B transition using two real current Hacker News events. The changed event forced a fresh briefing before reuse resumed; naturally observed rank transition is not claimed.
Controlled first-party benchmark · baseline quality TOOL_NECESSITY_PASS_BEST_BASELINE_UNPROVEN · semantic PASS · repeatability 3/3
Integration path →
75%
Shared generated image artifact
4 → 1 provider executions
≥$0.212 → ≥$0.053 · avoided ≥$0.159
Only valid when all callers explicitly accept one shared generated artifact and diversity is not required.
Controlled first-party benchmark · baseline quality CONTRACT_DEPENDENT · semantic PASS · repeatability 3/3
Integration path →