Base Qwen
24 pairs50.0% 0.0%
Hallucinating responses 100% reduction
- Exact success
- 12.5% 100%
- Induced errors
- 0.0%
- Producer run
- 40 steps
- Boundary
- Found
FUROS builds systems for planning, coordinating, and executing complex operations at scale.
Plan. Coordinate. Execute.
Reduced.2
A First Generation AI Hallucination Reduction Engine.2
Hallucinating responses by domain
All available paired runs
reasoning_effort: none; Qwen used its instruct configuration without a reasoning setting. Every recorded reasoning-token count is 0.external_holdout inside the benchmark suite, but its authorship and isolation have not been independently verified. No production or independently validated result is claimed.Golden Gate / AHT-Bench
Latest boundary-triggered paired run per model and domain 156 response pairs
15 / 15 / 18 / 18 / 24 paired responses
15 / 15 / 18 / 18 / N/A paired responses
Latest Antares run after the LLM auditor was disabled
100% relative reduction
+87.5 percentage points
+100 percentage points
100% relative reduction
Latest Antares runs after the auditor change
50.0% 0.0%
Hallucinating responses 100% reduction
No boundary
The producer completed all 40 Antares steps without a machine-verifiable hallucination.
Latest external-holdout run status
Boundary found
12 0
Hallucinating responsesNo boundary
40 steps
No paired comparison generatedScorer 3.2
Gate-only
LLM auditor not called