harveyai / harvey-labs Benchmark

Agent capabilities for legal work

Legal work is where nearly right is wrong.

harvey-labs is a benchmark built to evaluate and improve agent capabilities for supporting legal work. The reason it exists is that a plausible answer and a defensible one look identical until someone checks.

01

Evaluate

Give the same task to different agents and get a number back, so a claim about capability has something behind it.

02

Improve

The project puts improvement in its own description, which means the score is meant to feed work rather than end an argument.

03

Supporting, stated plainly

The description says supporting legal work. That word is doing real scoping, and a design should keep it rather than round it up.

harveyai/harvey-labs Design study by Goose. Not affiliated with the project.