Expert data for frontier systems

Models have read everything. They've practiced nothing.

Lydgate builds the practice. Expert-authored tasks, verifiable environments, private benchmarks — and the platform that produces and grades them. Written by people who do the work for a living, checked by machines that can prove it was done right.

Finance & marketsEngineering & softwareQuantitative research
Task registry Representative sample · not live data
IDDomainTaskAuthorGraderVerdict
LYD-4471FinanceFive-year LBO with debt schedule and cash sweepVP, sponsors coveragecell-level diffVerified
LYD-4472EngineeringReproduce and patch a flaky integration testStaff engineer, distributed systemssuite × 200 seedsVerified
LYD-4473QuantClosed-form price under stochastic volatilityPhD, mathematical financesymbolic equivalenceVerified
LYD-4474FinanceRestate segment revenue from the 10-K footnotesAssociate, equity researchsource-citation checkRejected
LYD-4475EngineeringMigrate a service across an auth boundaryPrincipal engineer, platformharness, 40 assertionsIn review
LYD-4474 failed on an unsourced figure in row 14.Every task ships with the grader that verifies it. Work that fails review does not ship.
What we build

Four things labs cannot buy off the shelf

Scraped text taught models to talk. It did not teach them to close a deal, ship a fix, or defend a number. That signal only exists where the work is actually done — and it needs a machine to collect it honestly.

Training data

Expert demonstrations

Supervised trajectories and preference pairs authored by credentialed practitioners — with the reasoning that produced them, not just the answer.

  • SFT demonstrations
  • Preference & comparison sets
  • Rubric-scored reasoning traces
  • Adversarial rewrites
Environments

Verifiable RL tasks

Executable environments with real tools and programmatic reward. If the model does the job, the grader proves it. If it doesn't, the grader says why.

  • Tool-use & agentic workflows
  • Deterministic reward functions
  • Multi-step, long-horizon tasks
  • Partial-credit rubrics
Evaluation

Private benchmarks

Held-out suites built from work product that has never been on the public internet — so a score means capability, not contamination.

  • Contamination-resistant by construction
  • Bespoke capability targets
  • Human expert baselines
  • Versioned, re-runnable
Platform

The machine that makes it

Data this specific can't be produced by a spreadsheet and a contractor pool. We built the workbench, the grader runtime, and the review system — and we'll run it in your cloud if you'd rather own it.

  • Expert authoring workbench
  • Sandboxed grader runtime
  • Review & adjudication queues
  • Provenance, versioning, delivery
See how the platform works
Two sides, one standard

We acquire the expertise. You get the signal.

Lydgate runs both halves of the problem: recruiting and vetting the practitioners who hold rare knowledge, and shaping what they produce into data a training run can actually use.

For AI teams

Data built to a spec you write

Tell us the capability you're missing. We assemble the practitioners, design the tasks, build the graders, and deliver against acceptance criteria you set in advance.

Engagement
Scoped pilot, then standing capacity
Delivery
Your format, your schema, your infrastructure
Ownership
Exclusive by default — we don't resell your data
How engagements work
For experts

Get paid for what you already know

Bankers, engineers, quants and researchers write the tasks that teach frontier models their field. Remote, project-based, on your own hours.

Who qualifies
Practicing professionals with verifiable experience
Commitment
Project-based — most contribute alongside a full-time role
Standard
Every submission is reviewed. Quality sets your rate.
What the work looks like
Method

Nothing ships that a machine can't check

Every piece of work moves through the same four gates, enforced by the platform rather than by good intentions. The order matters: we build the grader before we commission the task, so nobody is guessing what "correct" means.

Gate 01

Specify

We write the capability target and the acceptance criteria with your team, in plain language, before anyone is hired.

Gate 02

Build the grader

The reward function or rubric comes first. If correctness can't be checked programmatically, we redesign the task until it can.

Gate 03

Commission

Vetted practitioners produce the work in the tools they use professionally — spreadsheets, repositories, notebooks, filings.

Gate 04

Adjudicate

The grader runs, a second expert reviews, and disagreements go to a senior adjudicator. Rejections come back with reasons.

Read the method in full
Domains

Deep where it's hard to go

We go narrow on purpose. Depth in a domain is what makes the data worth buying — and what makes the graders honest.

Finance & markets

Work product from people who build the models that move capital — where a wrong number has consequences and every figure traces back to a filing.

M&A / LBOEquity researchCreditFP&APrivate markets

Engineering, software & math

Long-horizon technical work with unambiguous outcomes — code that compiles and passes, proofs that check, systems that stay up under load.

Distributed systemsCompilersFormal proofData scienceHardware
“…to pierce the obscurity of those minute processes.”
George Eliot · Middlemarch · 1871
The name

Tertius Lydgate is Eliot's country doctor: a man convinced that the answers were hiding in the details nobody had bothered to record, and that method beats reputation.

That's the whole thesis. The knowledge that makes an expert an expert was never written down for the internet to scrape. It lives in the minute processes — the intermediate step, the discarded approach, the reason the number was wrong. We record it, and we check it.

More about Lydgate

Tell us what your model can't do yet

We'll tell you whether it's a data problem, and what it would take to fix it. One inbox, read by the people doing the work.