Ragnarok Solutions / Red team We red team AI

Find where your AI is wrong before someone else does.

We find failures on purpose, against a written standard, and prove each one. You get the wrong outputs, the reason each is wrong, the check that would have caught it, and what to fix first.

The ten-day review

Offered now

You give us a sample of what your AI has produced and read access to the records it drew on. We re-derive every result, rule first, and hand you a written report. Nothing in your systems is changed.

  1. 01Up to 25 outputs sampled across your highest-stakes uses
  2. 02Each re-derived from source with the rule stated before the figure
  3. 03A defect register with the cause and the missing check for each
  4. 04A one-hour readout with your decision makers
Fixed fee

Quoted before we start

Ten working days once access and the sample are agreed. About three hours of your time.

  • If we find nothing wrong, the fee is the same and you hold written evidence that your AI held up on a defined sample.
  • The full fee is credited against any build we do for you afterward.

How the work is checked

the same method on every job
Days 1 to 2

Scope and sample

A one-hour call. We agree which outputs matter most, how we will read your records, and what "right" means for each, in writing, before we compute anything.

Days 3 to 8

Re-derive and check

Every output is derived twice, by separate seats in our pipeline that cannot see each other's result. Disagreements are resolved on evidence. A closing check reads our own draft before you see it.

Days 9 to 10

Report and readout

The defect register, the evidence behind each entry, the missing checks ranked by exposure, and an hour with your decision makers to walk it.

Everything on this line

ask for any of it
Offered now

Grader and reward-hack review

For environment vendors, labs and benchmark owners. We attack the grader before the model does: no-op and junk submissions, example copying, keyword stuffing, constants that hit score floors, fixture leaks. You get each exploit, proven, and the hardened grader.

On request

Agent and workflow red teaming

For any company that has put an agent into a real workflow. We test the whole chain, not only the model: tools, hand-offs, gates and what happens when a step fails silently.

On request

Prompt-injection and tool-abuse testing

Agents that read vendor PDFs and email can be steered by what they read. We test yours with documents built to do that, in your trade's own paperwork.

On request

Test a vendor's AI before you buy it

A demo is not a test. We run the product through a test company it has never seen and hand you a scorecard before you sign. For owners, controllers, sponsors and public buyers.

On request

Model bake-off

Two or more models or vendors, the same tasks from your own work, one fixed fee, one table of results.

On request

Domain benchmarks

Test sets in trades the public benchmarks do not cover: upstream oil and gas, structural analysis, construction compliance, anatomy and biomechanics.

What waiting costs

The error is already in a report. The only open question is who finds it.

If it is us, you get a private list and the fix. If it is a lender, a working-interest partner, a prime contractor or a regulator, you get a conversation you did not choose. A review costs a fixed fee and three hours of your time.

Questions we get

Is our data safe with you?

We work read only, inside your accounts where possible, under a written agreement. We take the sample and the records it depends on, nothing else, and return or destroy them at the end. The report is yours.

You also build AI. Can you test your own work?

We do, and we tell you when we are doing it. You are always free to have a third party repeat any test of a system we built. We will hand them the test set.

Which kinds of AI qualify?

Anything that produces a number, a document or a decision from your records: chat assistants your staff use, automated pipelines, a vendor product with a model inside it.

Can you just fix it instead of reviewing it?

We can, and the review fee is credited against the build. We still start with the review, because a fix that has not measured the problem is a guess.

Start smaller. Send us one output.

One number your AI produced, with the records it came from. We check it free and show you the evidence.