Scope and sample
A one-hour call. We agree which outputs matter most, how we will read your records, and what "right" means for each, in writing, before we compute anything.
We find failures on purpose, against a written standard, and prove each one. You get the wrong outputs, the reason each is wrong, the check that would have caught it, and what to fix first.
You give us a sample of what your AI has produced and read access to the records it drew on. We re-derive every result, rule first, and hand you a written report. Nothing in your systems is changed.
Quoted before we start
Ten working days once access and the sample are agreed. About three hours of your time.
A one-hour call. We agree which outputs matter most, how we will read your records, and what "right" means for each, in writing, before we compute anything.
Every output is derived twice, by separate seats in our pipeline that cannot see each other's result. Disagreements are resolved on evidence. A closing check reads our own draft before you see it.
The defect register, the evidence behind each entry, the missing checks ranked by exposure, and an hour with your decision makers to walk it.
For environment vendors, labs and benchmark owners. We attack the grader before the model does: no-op and junk submissions, example copying, keyword stuffing, constants that hit score floors, fixture leaks. You get each exploit, proven, and the hardened grader.
For any company that has put an agent into a real workflow. We test the whole chain, not only the model: tools, hand-offs, gates and what happens when a step fails silently.
Agents that read vendor PDFs and email can be steered by what they read. We test yours with documents built to do that, in your trade's own paperwork.
A demo is not a test. We run the product through a test company it has never seen and hand you a scorecard before you sign. For owners, controllers, sponsors and public buyers.
Two or more models or vendors, the same tasks from your own work, one fixed fee, one table of results.
Test sets in trades the public benchmarks do not cover: upstream oil and gas, structural analysis, construction compliance, anatomy and biomechanics.
The error is already in a report. The only open question is who finds it.
If it is us, you get a private list and the fix. If it is a lender, a working-interest partner, a prime contractor or a regulator, you get a conversation you did not choose. A review costs a fixed fee and three hours of your time.
We work read only, inside your accounts where possible, under a written agreement. We take the sample and the records it depends on, nothing else, and return or destroy them at the end. The report is yours.
We do, and we tell you when we are doing it. You are always free to have a third party repeat any test of a system we built. We will hand them the test set.
Anything that produces a number, a document or a decision from your records: chat assistants your staff use, automated pipelines, a vendor product with a model inside it.
We can, and the review fee is credited against the build. We still start with the review, because a fix that has not measured the problem is a guess.
Start smaller. Send us one output.
One number your AI produced, with the records it came from. We check it free and show you the evidence.