Overview · Domains & Benchmarks

Boundless Intuition

The AI you can Trust

Boundless Intuition builds the verification infrastructure for AI. We turn a company's policies, regulations, and domain's expert knowledge into machine-checkable logic, then verify every high-stakes AI action/output before it is trusted or executed.

When an AI agent tries to perform a consequential action, such as approving a tax filing, calculating a clinical dosage, authorizing a payment, or changing a firewall rule.

Our formal verification engine either produces a mathematical proof that the action satisfies the required rules or blocks it. Instead of asking companies to trust AI because it sounds confident, we give them machine-checkable proof that its decisions are correct, auditable, and compliant.

Our long-term vision is to become the foundational infrastructure for verified intelligence in high-stakes / mission-critical domains.

100%
Verified accuracy on RuleArena aviation
14×
Lower inference cost on aviation
100%
Verified accuracy on MedCalc-Bench
8.5×
Lower cost per correct answer on tax

Domains & Benchmarks

01Aviation Verification

RuleArena

We evaluate whether separating semantic extraction from deterministic execution improves correctness on RuleArena, an open evaluation benchmark for rule-guided reasoning, using its airline baggage fee domain.

Our pipeline of formalizers, a checker and a formal layer, improves frontier-model accuracy from 54% and 61% to 100%, while reducing inference cost by approximately 14× and latency by a factor of 11×.

Accuracy by arm

Correct answers out of 100 RuleArena cases

Fig 1Accuracy of every arm on the 100 RuleArena airline cases. Verification lifts both frontier tiers to a perfect score, and carries the budget tier from 3% to 82%.
02Clinical Verification

MedCalc-Bench

Our pipeline increased raw accuracy from 61-79% across all four Claude tiers to 100% at 29 ms at marginal cost per question.

Our audit even found bugs in MedCalc-Bench's own ground truth (wrong methadone conversion factors; 5 of 12 "impossible" questions have debatable labels).

Fable 5 (the most expensive model) scored below Sonnet and Opus; resampling flips only 11.7% of wrong answers. Errors are systematic.

Fig 2Baseline against verified across seven clinical metrics. The verified arm reaches the ceiling on every axis; the baseline gives ground on sensitivity, mimic accuracy, and run-to-run consistency.
03Statutory Verification

Tax code (Catala kernels, FR + US statutes)

We improved the performance of the non-frontier & open source models equivalent to that of frontier models with our pipeline.

Easy-case corpus

Our verification pipeline enables a cheap, non-frontier model to match frontier-model accuracy at 8.5× lower cost per correct answer, while reducing output tokens by 10 to 16×.

Hard-case French corpus

On the hard-case French corpus, frontier-model accuracy drops to 83.7%, while cheap non-frontier models fall to 45.6%. Our verification pipeline restores frontier-model accuracy to 100% and, with iterative verification, also brings the cheap non-frontier model to 100%, achieving this at 4× lower cost than the frontier-model baseline.

Cost against accuracy

Cost per correct answer (USD, log scale)

Cheap · verified
98.9% · $0.00094
Cheap · verified + loop
100% · $0.0044
Frontier · verified
100% · $0.0059
Cheap · baseline
62% · $0.0119
Frontier · baseline
98.9% · $0.0185
Frontier · verified + loop
100% · $0.0200
Fig 3Cost per correct answer against accuracy. Verification (blue) and iterative verification (green) dominate the unaided baselines (red): every verified arm is both cheaper and more accurate than the frontier baseline.
Ongoing research

Domains in progress

The three domains above are the ones we have taken far enough to report. The same pipeline is now being applied to other areas whose rules are already close to a formal specification - security control systems, where access and firewall policy is written as rules long before anyone writes code, and payment authorization, where every control has to be auditable regardless.

That work is earlier than what is on this page, so we are not putting numbers to it yet. A domain gets published when it clears the same bar as the three above: a kernel checked by hand against the source rules, a full run across the benchmark, and every figure generated from that run's logged data. Early results are promising, and we will keep adding them here as they land.

Team Boundless IntuitionJul 30, 2026