Onboarding ↗

Rubric Writing Guidelines

Start here

An AI agent can now be handed a finance brief and a data room and come back with the deck and the model. This project measures whether that work is actually good — not good for a machine; good by the standard of the people who do it for a living. You are one of those people, and this is how your judgement gets written down.

The project

A case starts in Phase 1: practitioners build it end to end — the brief, the data room, and a finished deck and model at IC standard, called the golden. In Phase 2, this phase, the case becomes a test. The brief becomes prompts an AI agent attempts, and experts write the rubrics those attempts are graded against. The graded results go to the teams building the agents: it is how the agents are measured, and what pushes them to improve.

Your part

You take a case whose golden is already built and checked, and write the rubric for each sub-task: five to twenty criteria that pin down, in advance and in writing, what a good deck or model must contain. You grade nothing by hand. An AI judge applies your rubric to every attempt, one criterion at a time.

Why it matters

The rubric is the only place your expertise enters the system, and it grades at a scale no reviewer could — every attempt, every run, long after the day you wrote it. Whatever your criteria let pass is what counts as good from then on. A sharp rubric puts a practitioner's judgement inside every grading run; a loose one grades noise — and noise is worse than nothing, because it looks like measurement.

New to AI work?

Expected — most people arriving here are. Nothing below assumes an AI background: the ten sections are the whole method, and the six terms that open §1 are the whole vocabulary. The one thing nobody can supply for you is the thing you already have — knowing what good work looks like. These guidelines exist to get it out of your head and into rows a machine can apply.

The standard, one section at a time. Search it, run the checker on your own rows, tick off the gates. From the one-page quick reference:

5–20rows per sub-task
5Criticals, minimum
1 in 5Criticals a strong attempt should fail
90%inter-judge agreement, three runs
4–5numbers per row, maximum

The ten sections

Three questions that decide most rows

Never

First time? Work through the onboarding module ↗ — eight exercises, one per idea.