Boundless Intuition
The AI you can Trust
Boundless Intuition builds the verification infrastructure for AI. We turn a company's policies, regulations, and domain's expert knowledge into machine-checkable logic, then verify every high-stakes AI action/output before it is trusted or executed.
When an AI agent tries to perform a consequential action, such as approving a tax filing, calculating a clinical dosage, authorizing a payment, or changing a firewall rule.
Our formal verification engine either produces a mathematical proof that the action satisfies the required rules or blocks it. Instead of asking companies to trust AI because it sounds confident, we give them machine-checkable proof that its decisions are correct, auditable, and compliant.
Our long-term vision is to become the foundational infrastructure for verified intelligence in high-stakes / mission-critical domains.
Domains & Benchmarks
RuleArena
We evaluate whether separating semantic extraction from deterministic execution improves correctness on RuleArena, an open evaluation benchmark for rule-guided reasoning, using its airline baggage fee domain.
Our pipeline of formalizers, a checker and a formal layer, improves frontier-model accuracy from 54% and 61% to 100%, while reducing inference cost by approximately 14× and latency by a factor of 11×.
Accuracy by arm
Correct answers out of 100 RuleArena cases
MedCalc-Bench
Our pipeline increased raw accuracy from 61-79% across all four Claude tiers to 100% at 29 ms at marginal cost per question.
Our audit even found bugs in MedCalc-Bench's own ground truth (wrong methadone conversion factors; 5 of 12 "impossible" questions have debatable labels).
Fable 5 (the most expensive model) scored below Sonnet and Opus; resampling flips only 11.7% of wrong answers. Errors are systematic.
Tax code (Catala kernels, FR + US statutes)
We improved the performance of the non-frontier & open source models equivalent to that of frontier models with our pipeline.
Our verification pipeline enables a cheap, non-frontier model to match frontier-model accuracy at 8.5× lower cost per correct answer, while reducing output tokens by 10 to 16×.
On the hard-case French corpus, frontier-model accuracy drops to 83.7%, while cheap non-frontier models fall to 45.6%. Our verification pipeline restores frontier-model accuracy to 100% and, with iterative verification, also brings the cheap non-frontier model to 100%, achieving this at 4× lower cost than the frontier-model baseline.
Cost against accuracy
Cost per correct answer (USD, log scale)
- Cheap · verified
- 98.9% · $0.00094
- Cheap · verified + loop
- 100% · $0.0044
- Frontier · verified
- 100% · $0.0059
- Cheap · baseline
- 62% · $0.0119
- Frontier · baseline
- 98.9% · $0.0185
- Frontier · verified + loop
- 100% · $0.0200
Domains in progress
The three domains above are the ones we have taken far enough to report. The same pipeline is now being applied to other areas whose rules are already close to a formal specification - security control systems, where access and firewall policy is written as rules long before anyone writes code, and payment authorization, where every control has to be auditable regardless.
That work is earlier than what is on this page, so we are not putting numbers to it yet. A domain gets published when it clears the same bar as the three above: a kernel checked by hand against the source rules, a full run across the benchmark, and every figure generated from that run's logged data. Early results are promising, and we will keep adding them here as they land.

