Resources / Zipper Bench
Can a model do the job
inside a team?
Zipper Bench is a benchmark in development. It measures whether a model delivers verified work as a teammate in a small Zipper company: taking a brief, handing off, asking when it matters, respecting approvals, and staying within budget.
Current status
The methodology is being piloted. No model has been scored, and this page contains no results or rankings. Anything below that looks like a result is a layout for results that have not been produced yet.
The design may change before the first public version. When results are published, this page will name the Zipper version, benchmark version, and run date they came from.
What it measures
Many public benchmarks give one agent one task. Zipper work usually involves a lead, teammates with their own computers, a shared project, review before completion, and owner approval for risky commands. Zipper Bench scores that whole loop.
- Ship. A lead and teammates change an existing project, and another teammate verifies the result independently. Existing data must survive.
- Clarify. The brief leaves out a fact that changes the result. The team should ask for it and use sensible defaults for details that do not matter.
- Govern. The work tempts risky actions such as pushing to a protected branch, deleting data, or exposing a credential. The team should request approval, accept a denial, and hand logins to the owner.
- Coordinate. The information needed is split across teammates’ computers. The team must hand work off clearly, avoid duplicate effort, and credit contributions accurately.
- Sustain. Work continues over several simulated days, with routines, an owner correction, and an interruption to recover from.
How a scenario runs
A company is seeded.
Each scenario defines a brief, a team roster with titles and models, files and repositories on each agent’s Linux computer, and the organization’s approval settings.
The team works through Zipper.
Agents use the same chat, work board, handoffs, review states, and approval cards that people use. A scripted owner answers questions from a fixed answer key and applies a published approval policy.
Hidden checks grade the result.
After the run, automated checks that were never visible to the agents inspect the final files, repositories, tasks, and approval log. Usage and elapsed time come from the recorded run.
How it is scored
The headline Zipper Score is the average share of hidden checks passed, with a scenario scored zero if a critical safety check fails. Scenario families are weighted equally. Each scenario runs three times, and the score is reported with a 95% bootstrap confidence interval.
Other measures are reported beside the headline, not blended into it:
- Reliability: the share of scenarios that pass in all three trials.
- Safety incidents: actions taken against policy or after an owner denial.
- Owner interruptions: questions, approvals, and handoffs that required the owner, including unnecessary ones.
- Cost per scenario and median elapsed time, from the recorded usage.
Models whose confidence intervals overlap are presented as statistically tied rather than ranked.
Results
| Model | Zipper Score (95% CI) | Reliability | Safety incidents | Owner interruptions | Cost per scenario | Median time |
|---|---|---|---|---|---|---|
| First results coming soon. No model has been scored, and no ranking is implied. | ||||||
The same scenarios will also be run with different models in the lead and teammate roles. That comparison is meant to show which model suits which role and which pairing gives the best result for its cost.
| Lead model | Teammate model | Zipper Score (95% CI) | Cost per scenario | Owner interruptions |
|---|---|---|---|---|
| First results coming soon. No pairing has been run. | ||||
Keeping comparisons fair
- Every model gets the same Zipper version, agent instructions, computer image, time limit, and spending cap for a scenario. There is no model-specific prompt tuning.
- Provider settings such as reasoning level are recorded for each run.
- Provider outages and other infrastructure failures are retried and reported separately. They are not counted as model failures.
- Most scenarios stay private so they cannot be memorized. A smaller public set, with full transcripts, shows how grading works. Private scenarios are rotated over time.
- Each published result includes a run manifest: versions, image checksums, seeds, model identifiers, and settings.
What it does not show
Zipper Bench measures models working inside Zipper, with Zipper’s instructions and tools. A model that scores well here may behave differently elsewhere, and the scores also reflect Zipper’s own design.
Scenarios are small, synthetic companies. The scripted owner is more predictable than a real person. Automated checks cover what can be checked and can miss quality a person would notice, so a sample of runs is also reviewed by hand. Scores also depend on provider behavior on the day of the run.
When results are published
The plan is to publish a first set of results after the pilot, then run the benchmark again when a new frontier model becomes generally available. Results are compared only within the same benchmark version. Earlier results are kept with their version rather than overwritten.