ResearchCompact decision models: where we’re headed →

Verified fast paths for repeated agent work

Midbar runs inside the agents your team already uses. When a request matches a procedure your team has done before, Midbar runs a compiled, checked version of it on your machine. Anything new or unclear goes back to the agent.

The bet

Your agents keep re-deriving work your team already knows.

Coding agents are good at new problems. On the tenth identical request this week they still start from scratch: rediscover the table, re-guess the house convention, re-run the same reasoning, and sometimes land somewhere slightly different.

Your team’s history already contains those procedures: which system, which fields, how values are stored, which checks prove it worked. Midbar turns the procedures that repeat into small, reviewed programs.

That shrinks the model’s job. It no longer writes the change. It decides which known operation was asked for and fills in its parameters, and the program does the rest, checked before anything is proposed.

Compile what your team already knows. Let the model only decide what was asked.

How it works

Repeated work in. Verified fast path out.

The model never writes the change itself. It chooses the skill and fills its parameters; the compiled skill does the work, and the checks decide whether it is proposed.

01

Capture

Record work sessions inside your environment, with tool results and the exact repository state, so a past task can be replayed and checked later.

02

Compile

Turn a procedure that keeps coming back into a typed skill: its parameters, your storage conventions, its guards, and the checks that prove it worked.

03

Decide

A small local model maps the request onto that skill. If two readings would do different things, Midbar asks. If nothing fits, it hands the request back to the agent.

04

Verify

Dry-run against current state, confirm that only the requested change happens, then propose it, or run it under the approval policy you set.

Compared to what you already have

Not another agent. A checked lane inside yours.

Agent alone
Agent + prompt notes
Agent + Midbar fast path
Repeated request
Plans it from scratch
Re-reads your notes, then plans
Your conventions
Guessed from context
Described in prose
Checking the result
Reads its own output
Reads its own output
When it’s unclear
Often picks one reading
Often picks one reading
Where it runs
Provider API
Provider API
What you keep
Chat history
Prompt files

We’re not replacing your agent. Click a Midbar cell for a short note.

Deployment

Inside the agents you use. On machines you control.

Midbar plugs into Cursor, Claude Code and other MCP-capable agents as a local server and a command-line tool. Capture, skills and checks run where your work already happens, so known requests don’t have to leave the building.

  • Nothing to migrate

    Your agent stays the front door. Midbar takes the requests it can verify and hands the rest straight back.

  • Data stays put

    Sessions, compiled skills and decision records live on your machines, not in someone else’s training set.

  • Small where it counts

    The goal is a decision step for known work that runs on ordinary local hardware. We’re measuring that now, not assuming it.

Evaluation

Score the effect, not the text

A convincing answer can still change the wrong field. We judge a run by what it actually did to the system, checked independently of the model that produced it.

Did the right thing change?

The end state matches the request, with exact types: true is not 1, and a string is not a number.

Did anything else change?

Every field and row the request didn’t mention stays exactly as it was.

Did it know when to stop?

A clear question or an honest hand-back beats a confident wrong change.

Would it hold up on fresh work?

Measured on new requests it was never tuned on, not on the examples it learned from.

  • Expert path
  • Midbar (checked path)
  • Generic + human fixes (✕)
Illustration, not a measurement. A run is judged on the end state it produces, not on how convincing the output looked.

Illustrative animation of one request, not measured data. The expert and Midbar complete the same checked steps. An unchecked path picks a wrong tool, skips a step, and needs human corrections marked with X.

Research direction

Compact decision models

On a repeated request, the hard part is rarely writing the change. It’s deciding what was meant. We’re researching a compact model, tens of millions of parameters rather than billions, whose only job is to score which known operation a request asks for.

It keeps competing readings alive instead of collapsing to one guess, and lets Midbar propose only when every remaining reading would do the same thing. Otherwise it asks, or hands the request back to your agent.

Read the research note →

  1. Built

    The decision layer, with no model in the loop

    A typed effect language, a reference interpreter, and a decision step that proposes, asks, abstains or declines. Tested on hand-written cases, including the ones designed to trick it.

  2. In progress

    Real requests, labelled by the people who made them

    We’re collecting fresh operator requests and marking what each one was meant to do, to measure how much real language the effect language can express.

We’ll publish results with their sample sizes and caveats. Until then, this is a direction, not a benchmark claim.

Good fits

Work that repeats. Work that has to be exactly right.

Operations

  • Configuration changes with house conventions
  • Runbooks that touch internal systems
  • Checks that look the same every day
  • Changes that must not touch anything else

Engineering

  • Repo-specific migrations and refactors
  • Release and deploy steps
  • First-response incident playbooks
  • Fixture and test upkeep after interface changes

Approval-heavy work

  • Changes a reviewer must sign off
  • Routing against your own rules
  • Requests where “close enough” is wrong
  • Anything you want a decision record for

Bring one workflow your team repeats every week

We’ll show you which parts can run as a verified fast path, and which should stay with your agent.