Capture
Record work sessions inside your environment, with tool results and the exact repository state, so a past task can be replayed and checked later.
Midbar runs inside the agents your team already uses. When a request matches a procedure your team has done before, Midbar runs a compiled, checked version of it on your machine. Anything new or unclear goes back to the agent.
The bet
Coding agents are good at new problems. On the tenth identical request this week they still start from scratch: rediscover the table, re-guess the house convention, re-run the same reasoning, and sometimes land somewhere slightly different.
Your team’s history already contains those procedures: which system, which fields, how values are stored, which checks prove it worked. Midbar turns the procedures that repeat into small, reviewed programs.
That shrinks the model’s job. It no longer writes the change. It decides which known operation was asked for and fills in its parameters, and the program does the rest, checked before anything is proposed.
Compile what your team already knows. Let the model only decide what was asked.
How it works
The model never writes the change itself. It chooses the skill and fills its parameters; the compiled skill does the work, and the checks decide whether it is proposed.
01
Record work sessions inside your environment, with tool results and the exact repository state, so a past task can be replayed and checked later.
02
Turn a procedure that keeps coming back into a typed skill: its parameters, your storage conventions, its guards, and the checks that prove it worked.
03
A small local model maps the request onto that skill. If two readings would do different things, Midbar asks. If nothing fits, it hands the request back to the agent.
04
Dry-run against current state, confirm that only the requested change happens, then propose it, or run it under the approval policy you set.
Compared to what you already have
We’re not replacing your agent. Click a Midbar cell for a short note.
Deployment
Midbar plugs into Cursor, Claude Code and other MCP-capable agents as a local server and a command-line tool. Capture, skills and checks run where your work already happens, so known requests don’t have to leave the building.
Your agent stays the front door. Midbar takes the requests it can verify and hands the rest straight back.
Sessions, compiled skills and decision records live on your machines, not in someone else’s training set.
The goal is a decision step for known work that runs on ordinary local hardware. We’re measuring that now, not assuming it.
Evaluation
A convincing answer can still change the wrong field. We judge a run by what it actually did to the system, checked independently of the model that produced it.
The end state matches the request, with exact types: true is not 1, and a string is not a number.
Every field and row the request didn’t mention stays exactly as it was.
A clear question or an honest hand-back beats a confident wrong change.
Measured on new requests it was never tuned on, not on the examples it learned from.
Illustrative animation of one request, not measured data. The expert and Midbar complete the same checked steps. An unchecked path picks a wrong tool, skips a step, and needs human corrections marked with X.
Research direction
On a repeated request, the hard part is rarely writing the change. It’s deciding what was meant. We’re researching a compact model, tens of millions of parameters rather than billions, whose only job is to score which known operation a request asks for.
It keeps competing readings alive instead of collapsing to one guess, and lets Midbar propose only when every remaining reading would do the same thing. Otherwise it asks, or hands the request back to your agent.
Built
A typed effect language, a reference interpreter, and a decision step that proposes, asks, abstains or declines. Tested on hand-written cases, including the ones designed to trick it.
In progress
We’re collecting fresh operator requests and marking what each one was meant to do, to measure how much real language the effect language can express.
Not yet
The small encoder hasn’t been trained. We haven’t measured its accuracy, latency or cost on real work, so we don’t quote numbers.
We’ll publish results with their sample sizes and caveats. Until then, this is a direction, not a benchmark claim.
Good fits
Operations
Engineering
Approval-heavy work
Notes
Start with the thesis for the full argument. These notes go deeper on compiled skills, verification, data, and where our research is headed.
Research
A small model that decides what a request means, keeps competing readings, and proposes only when they agree.
Read note →Systems
Encode what your team already knows in a reviewed program. The model only picks the operation and fills its parameters.
Read note →Evaluation
A convincing answer can still change the wrong field. Judge a run by what it did to the system.
Read note →Data
Work history shows which systems, fields and checks a task really needs, if it records results and state.
Read note →We’ll show you which parts can run as a verified fast path, and which should stay with your agent.