I paid frontier models to build a factory. Now that factory works while I sleep.
Overnight, my local AI lab researches new models, watches repositories, evaluates releases, drafts articles, checks code, transcribes media, and monitors security signals. It uses idle compute without asking me to babysit a terminal. By morning, the useful work has converged into one digest and one review queue.
I wake up to decisions, not chores. That changes the shape of my day.
The lab has graduated from a collection of clever scripts into an always-on operation. It can observe, reason, retrieve, draft, test, score, schedule, and propose work 24 hours a day. I keep authority over publishing, code merges, pushes, deletions, and semantic changes, but the daily cognitive grind keeps moving without me.
Codex built that operating model during a 9-hour-and-51-minute goal. I paid for frontier cloud reasoning to design it, attack it, implement it, and verify it. That cloud spend was worth every dollar. It created a recurring local operation that now runs on hardware I already own.
I paid the cloud to build the factory
I did not hand the entire job to one model and let it plan, build, grade, and approve itself. That is how one confident mistake becomes an entire system.
Codex split the goal into distinct seats. A root integrator owned the objective and the final claim. Claude Fable 5 planned the work. GPT-5.6 Sol challenged the plan and reviewed the evidence. GPT-5.6 Luna executed bounded slices. The reviewer did not implement the work it judged, and the executors could not widen their own assignments.
No model is unbiased. Separation of duties makes bias collide with another model before it reaches production. The planner can be wrong without implementing its mistake. The executor can miss something without grading its own work. The reviewer can reject an artifact without gaining the authority to rewrite history. The root still has to reconcile every result against the original goal.
That is what I bought with cloud inference. Frontier models acted as the construction crew. I paid them to build and inspect machinery that would keep operating after the session ended.
The economics change once that machinery moves local. In the last 24 hours, the lab completed 457 local-model calls. It processed 1.309M input tokens and 184k output tokens while avoiding an estimated $6.68 of cloud-equivalent cost. Across all tracked activity, it has completed 1,256 calls, processed 9.899M input tokens and 2.065M output tokens, and avoided about $60.68.
These figures estimate what the same workload would have cost in the cloud. The comparison prices local activity at Sonnet 5 rates of $3 per million input tokens and $15 per million output tokens. It also combines measured tokens with activity-based estimates. The number I care about is 457 pieces of model work in one day that did not create another cloud API bill.
I spent cloud money to build the factory. The factory now handles the recurring work at near-zero marginal API cost. That is an outstanding trade.
My night shift does not clock out
The lab has recurring lanes for the work that fractures my attention during the day.
It follows AI research and model releases. It watches repositories and changelogs. It evaluates local models against my workflows instead of trusting a public leaderboard. It drafts articles, replies, meeting preparation, and email intelligence. It checks documentation and code, transcribes recordings, analyzes media, monitors sites and security signals, and proposes repairs when it finds something worth pursuing.
A night-shift scheduler fills idle gaps across the machines. When I or a scheduled workflow needs the shared compute, it yields. When the hardware is free, it attacks the queue again. The lanes feed a morning digest and a review queue, so I make a small number of informed decisions instead of reconstructing a night of activity from logs.
Autonomous work while I sleep changes how I use every waking hour. The lab does not need me awake to make progress.
It can research a model release, test it, compare the output, draft a recommendation, and put that recommendation in front of me. It can find documentation drift, prepare a repair, and wait for approval. It can turn a recording into structured notes while another lane watches a repository and another evaluates a model.
The human gate stays. Publishing, semantic wiki edits, merges, pushes, deletions, and other consequential changes still require my approval. That boundary does not make the lab less autonomous. It gives the lab room to move fast without giving every model the keys to production.
The work continues around the clock. I keep the authority; the lab keeps the momentum.
Bonsai blew open the resource ceiling
Bonsai is a massive deal for this lab.
The lab's active large-model role moved from a roughly 19 GB Qwen3.6-35B-A3B package to a roughly 7.3 GB Bonsai 27B model-plus-projector package. One production change removed about 11.7 GB from the active large-model package. That is about 62 percent smaller.
On a local AI machine, 62 percent smaller is enormous. Every model competes for the same finite pool of memory. Cutting the active large-model package from 19 GB to 7.3 GB creates room for the always-on Hermes model, long-context work, scheduled evaluators, and concurrent workflows. More of the lab can stay alive at the same time.
Bonsai is a ternary low-bit model from PrismML based on Qwen3.6-27B. Its compact weights use three values: negative one, zero, and positive one. PrismML compressed and post-trained those weights. I did not retrain Bonsai. I tuned the production system around it: the serving profile, prompts, deterministic contracts, and evaluation harness that decide whether it can do useful work here.
We pushed it through structured extraction, tool use, vision, and long-context retrieval at several depths. Its production profile supports a 64K context window. We ran it beside the always-on Hermes workload. Every health sample passed during the coexistence soak, critical memory pressure never appeared, and swap did not grow.
This is not a claim that Bonsai became smarter than Qwen at everything. Historical task-faithful evaluation favored Qwen, and Bonsai's strict raw-judge result remains just below my promotion threshold. Qwen stays available as a rollback. Harder tasks can still escalate to a cloud model.
The win is bigger than a benchmark slogan. Bonsai produces contract-correct work for its assigned production role while consuming a fraction of the model footprint. It lets the whole lab do more work at once.
That smaller model package changed what one machine can carry. Bonsai earned a production role, freed capacity for the rest of the lab, and kept the stronger escape routes intact.
The breakthrough is the operating system
Model launches get the headlines, but the model alone did not create this lab. The breakthrough came from combining specialized local models, deterministic routing, independent evaluation, idle-time scheduling, and explicit authority boundaries.
Codex supplied the institutional structure. It decided who planned, who challenged, who executed, who verified, and who was allowed to change what. Bonsai changed the resource equation. Hermes gave the lab an always-on local reasoning seat. The night shift turned unused hours into productive hours. The review queue kept consequential decisions with me.
After 9 hours and 51 minutes of frontier-model orchestration, I had more than a repaired stack. I had a local AI operation that keeps working after I close the laptop, uses the machines I already own, and brings the consequential decisions back to me.
The cloud models built the factory. The local models now run the shift. By morning, I am looking at the part only I should decide.