claim. verify. sign.
An AI system's report of its own success is a self-signed certificate.
It proves the system said it, not that it happened.
Independent verification for AI systems that claim real-world outcomes. Turiya reruns the claimed action against ground truth and signs what actually happened, so anyone can check the result — including the people who won't take your word for it.
clickclaimed fixed / effect: fails
Here's exactly what happens.
One claim in, one signed verdict out. This is a real receipt, walked end to end.
Built for the people who can't grade their own homework.
Not software you install. A verification engagement.
One high-stakes workflow, re-run against ground truth inside your environment. Your code and data stay with you; what leaves is signed evidence you can pass to someone who won't take your word for it.
A verifier that never corrects itself is not a verifier. We publish our misses.
Full record, including every candidate we rejected and why. Read the methodology.
sha256:… $ awaiting verification…
claimed fixed / effect: fails
These numbers are signed. Verify them ↑
- Real verification, not a demo
- A signed receipt you keep
- Verifiable in your browser
- No call required
- Batch verification
- Z3-checked specs
- Assurance report
- Auditor-ready evidence
- Re-verification each quarter
- Drift detection
- AIUC-1 aligned evidence
- Named contact
- SSO / OIDC + RBAC
- Air-gap deployment
- SLA
- White-label for resellers
One price is public: the free pilot. Everything else is scoped, because a single number would misrepresent every engagement except one.
Any AI system that claims an effect. Cut by verifiability.
The engine is system-agnostic: predictive, generative, assistive, agentic, orchestrated, the whole tree. What decides whether we can verify a system is one question: does its claim leave a checkable effect?
checkable effect
what it did leaves a falsifiable trace: a flag, a code change, a grounded answer
we falsify it: re-run, re-check, sign
judgment-heavy
"correct" needs expert taste: was the loan fair, the diagnosis sound
signed expert judgment
unfalsifiable
no ground truth exists: open-ended creation, intent, alignment
we refuse: indeterminate, never a fabricated score
Signed receipts are published across 17 system types: code effects, fraud, credit, sanctions screening, fairness, anomaly detection, multi-agent orchestration, forecast, recommendation, computer vision, speech, scientific AI, RAG grounding, agents, truth, efficacy, and honesty. The rest reuse the same engine; we claim a receipt only once it exists.
The same verifier, against a real fraud model.
Run on the real ULB MLG dataset — 284,807 anonymized European card transactions, 0.17% fraud — the falsifier caught two real gaps: temporal drift and adversarial evasion.
XGBoost — the workhorse of production fraud detection — is 99.9% accurate on real transactions, yet its f1 still swings 1.59× across time, and gradient-guided perturbations evade 22.6% of the fraud the logistic model catches — 8.4% transferred to a black-box XGBoost. And the amount-evasion attack that fired on an earlier synthetic dataset doesn't fire on real data (0 of the fraud it caught was evaded) — the verifier reports that honestly rather than manufacturing a catch.
And on IEEE-CIS — 590,540 real e-commerce transactions — a model built on the naive interpretable signals (amount, card, address, email) is 94.4% accurate yet catches only 23.9% of fraud: the fraud is low-and-slow, structured to mimic legitimate behavior.
logistic regression 96.391% accurate evades 44/195 of caught fraud; the SAME adversarial samples transfer to a black-box XGBoost and evade 15/179 — mean applied L-inf perturbation 0.500
99.923% accurate; 0/179 of caught fraud evaded by halving amount
94.4% accurate yet catches only 23.9% of fraud — below the 96.5% 'always legitimate' baseline
f1 0.753 (recall 0.798) · f1 0.761 (recall 0.798) · f1 0.773 (recall 0.807) · f1 0.781 (recall 0.798)
f1 swings 0.538 to 0.857 across time segments (1.59x spread) — the validated number is a snapshot, not a property
5 signed fraud receipts, each Ed25519-signed over the record hash. Download any of them and re-verify with the public key.
One falsifier, seventeen system types.
Code, fraud, credit, sanctions screening, fairness, anomaly detection, multi-agent orchestration, forecast, recommendation, computer vision, speech, scientific AI, truth, efficacy, honesty, RAG grounding, multi-step agents — the same claim-vs-effect engine signs a receipt for each. Where the system under test was honest, the receipt says so. We don't manufacture a catch.
The newest catches — speech, scientific AI, and the XGBoost re-runs: a vosk recognizer's word error rate jumps from 20% to 90% under noise; XGBoost cannot extrapolate the CO2 trend (7.0 vs 1.2 ppm error); and fairness, forecast, and scientific were re-falsified against XGBoost — the findings hold. Read the signed receipt →
The fifth layer: the signed effect.
AI assurance has four crowded slices: monitoring (did the distribution shift?), security (can it be attacked?), capability (how good is it?), controls (are the processes right?). We verify the fifth: the effect (did it do what it claimed), and we sign the answer. The other four layers can complement the signed effect. We sign the ground-truth verdict. Not a benchmark score, not a PDF report — a signed receipt you can re-verify.
Independent verification is becoming a requirement.
EU AI Act enforcement powers took effect 2 August 2026. California signed SB 813 + AB 1405 in September 2026: an AI Auditor Registry that requires independent verification from parties who did not build the system. AIUC-1 certification mandates quarterly third-party re-testing. The question is no longer "should we verify?" It is "who signs the proof?"