Models have read everything. They've practiced nothing.
Lydgate builds the practice. Expert-authored tasks, verifiable environments, private benchmarks — and the platform that produces and grades them. Written by people who do the work for a living, checked by machines that can prove it was done right.
Four things labs cannot buy off the shelf
Scraped text taught models to talk. It did not teach them to close a deal, ship a fix, or defend a number. That signal only exists where the work is actually done — and it needs a machine to collect it honestly.
Expert demonstrations
Supervised trajectories and preference pairs authored by credentialed practitioners — with the reasoning that produced them, not just the answer.
- SFT demonstrations
- Preference & comparison sets
- Rubric-scored reasoning traces
- Adversarial rewrites
Verifiable RL tasks
Executable environments with real tools and programmatic reward. If the model does the job, the grader proves it. If it doesn't, the grader says why.
- Tool-use & agentic workflows
- Deterministic reward functions
- Multi-step, long-horizon tasks
- Partial-credit rubrics
Private benchmarks
Held-out suites built from work product that has never been on the public internet — so a score means capability, not contamination.
- Contamination-resistant by construction
- Bespoke capability targets
- Human expert baselines
- Versioned, re-runnable
The machine that makes it
Data this specific can't be produced by a spreadsheet and a contractor pool. We built the workbench, the grader runtime, and the review system — and we'll run it in your cloud if you'd rather own it.
- Expert authoring workbench
- Sandboxed grader runtime
- Review & adjudication queues
- Provenance, versioning, delivery
We acquire the expertise. You get the signal.
Lydgate runs both halves of the problem: recruiting and vetting the practitioners who hold rare knowledge, and shaping what they produce into data a training run can actually use.
Data built to a spec you write
Tell us the capability you're missing. We assemble the practitioners, design the tasks, build the graders, and deliver against acceptance criteria you set in advance.
- Engagement
- Scoped pilot, then standing capacity
- Delivery
- Your format, your schema, your infrastructure
- Ownership
- Exclusive by default — we don't resell your data
Get paid for what you already know
Bankers, engineers, quants and researchers write the tasks that teach frontier models their field. Remote, project-based, on your own hours.
- Who qualifies
- Practicing professionals with verifiable experience
- Commitment
- Project-based — most contribute alongside a full-time role
- Standard
- Every submission is reviewed. Quality sets your rate.
Nothing ships that a machine can't check
Every piece of work moves through the same four gates, enforced by the platform rather than by good intentions. The order matters: we build the grader before we commission the task, so nobody is guessing what "correct" means.
Specify
We write the capability target and the acceptance criteria with your team, in plain language, before anyone is hired.
Build the grader
The reward function or rubric comes first. If correctness can't be checked programmatically, we redesign the task until it can.
Commission
Vetted practitioners produce the work in the tools they use professionally — spreadsheets, repositories, notebooks, filings.
Adjudicate
The grader runs, a second expert reviews, and disagreements go to a senior adjudicator. Rejections come back with reasons.
Deep where it's hard to go
We go narrow on purpose. Depth in a domain is what makes the data worth buying — and what makes the graders honest.
Finance & markets
Work product from people who build the models that move capital — where a wrong number has consequences and every figure traces back to a filing.
Engineering, software & math
Long-horizon technical work with unambiguous outcomes — code that compiles and passes, proofs that check, systems that stay up under load.
“…to pierce the obscurity of those minute processes.”
Tertius Lydgate is Eliot's country doctor: a man convinced that the answers were hiding in the details nobody had bothered to record, and that method beats reputation.
That's the whole thesis. The knowledge that makes an expert an expert was never written down for the internet to scrape. It lives in the minute processes — the intermediate step, the discarded approach, the reason the number was wrong. We record it, and we check it.
More about Lydgate →Tell us what your model can't do yet
We'll tell you whether it's a data problem, and what it would take to fix it. One inbox, read by the people doing the work.