It cannot be memorised
The company is invented. The ground under it is real and refreshes every month: filings, grid and weather feeds, imagery, terrain. Last month's snapshot is the agent's world. This month's is the answer key.
For AI labs, environment vendors and model teams. We build the worlds, tasks and graders that models are trained and measured on, in physical industries where the people who know the work rarely write code.
The company is invented. The ground under it is real and refreshes every month: filings, grid and weather feeds, imagery, terrain. Last month's snapshot is the agent's world. This month's is the answer key.
Every task is scored by a computation, an open solver, or a record that arrives later on a statutory clock. No judge model has the last word on a number.
Scanned stubs, run tickets, plats, gauge sheets, logs, aerial and thermal images. An image is in a task only when the task cannot be solved without looking at it.
Before you deliver an environment, we try to beat its grader without doing the work. You get each exploit, proven, and a grader that holds. The fastest way to see how we work.
Deterministic graders and rubrics for one trade, written by people who have done the job. Each requirement graded once. Tolerances stated. No mirror criteria.
Tasks with difficulty measured against model performance before delivery. A determinism checklist and a manifest of every figure the task depends on ship with each one.
A full synthetic operating company for your sector, with its systems, its paper, its planted defects and its graders. Packaged for the open harness formats your team already runs.
Your ad hoc test scripts moved into a proper harness, with domain tasks added, run monthly as your gate.
A small vetted bench of field, construction and clinical experts, delivered with our reviewer pipeline as the quality gate. Not a marketplace.
A catalogue of living environments beyond energy, licensed non-exclusive or exclusive. Each sits on public data with an invented company on top, so we own what we license.
Domain results on public models from our private test sets, and entries in open arenas and environment hubs.
Nothing from a past programme is ever reused.
No task, rubric, grader or data from any client or platform programme appears in anything we build for you. We build the same kind of thing, new, from public sources and invented companies. That is a confidentiality rule, and it is also why we can license what we make.
A model will find the hole in your grader before your customer does. Then your customer will.
A gamed reward trains the wrong behaviour and the environment comes back. A review before delivery is a fixed fee and a few days.
Send us one grader. We will try to break it.
Tell us the domain and how the task is scored. We reply within one working day with how we would attack it and what a review would cover.