Enterprise AI adoption & deployment
I wake up to decisions, not chores.
A 24/7 local AI lab does the research, drafting, and checking overnight on hardware I own. The same operating discipline sits behind 40 enterprise GenAI deployments.
The lab, live.
While you sleep, this lab works. Seven stations run every night on hardware I own, each one smarter than the night before: the loop tests its own ideas and promotes only measured wins. Open a station for the measured proof.
-
Ingest
Ingest
4:30 AM, every night. The lab clocks in before I do.
-
Route
Route
The 64GB Mac Studio runs one big model at a time, nothing fighting for memory. A Mac mini serves one 9B that never swaps.
-
Run
Run
Email, meetings, code review, docs, scribe, publish, gmail digest: seven maintenance lanes running 24/7 on local models at zero marginal cost. Frontier models stay reserved for judgment.
-
Judge
Judge
Every draft the lab writes gets scored against my own style gate. One night of retraining against the gate's measured failures moved the writer bench from 88.8 to 97.0. Receipts, not vibes.
-
Learn
Learn
Judge scores feed a research loop that rewrites the lab's own prompts and rules. Promotion takes a measured win. Its first rule scored 98.0, then 94.0, so the loop said no and changed nothing. The rejection is the feature.
-
Gate
Gate
Fifty challenger models benched, 45 deleted. One 536-call preregistered experiment published INCONCLUSIVE. The gate answers to evidence, not excitement.
-
Ship
Ship
Every morning the lab reports to me: a narrated brief, drafts already judged, and a review queue holding the night's decisions for my approval. I wake to decisions, not chores.
Things I built and shipped.
-
https://github.com/jbelnick/cerebellum-local-ai-router
Local AI Router
Routes lower-risk AI work to local models with policy controls and a reviewable decision trail.
-
https://github.com/jbelnick/planbridge
PlanBridge
Local, read-only MCP connector that lets ChatGPT plan over an allowlisted workspace, then hands the frozen plan to Codex.
-
https://github.com/jbelnick/meeting-intelligence-pipelines
Meeting Intelligence Pipelines
Turns sanitized call notes into reviewed follow-up, risk flags, and named owners.
-
https://github.com/jbelnick/llm-judge-evals
Evaluation Harness
A golden-dataset judge that fails CI when output quality drifts.
-
https://github.com/jbelnick/meeting-intelligence-mcp
Meeting Intelligence MCP Server
Read-only meeting tools exposed behind a stable, packaged boundary.
Recent writing.
Practical notes on AI adoption and local models. All writing ->
-
Fitting Qwen3.8 Flash-Next on a 64 GB Mac Studio
How a 68 GB-resident model runs in 42 GB at better-than-reference speed with a row-granular NVMe loader for its 51B n-gram table
-
I Built a 24/7 Local AI Lab, Then Cut Its Biggest Model by 62%
I paid frontier models to build a 24/7 local AI factory, then cut its active large-model package by 62 percent and moved the recurring work onto hardware I already own.
-
The night my writer retrained twice
I retrained my local writer model twice in one evening, the second time on training data written against my style gate's measured failures, and its bench score went from 88.8 to 97.0. This post was drafted by the model it describes.
If you are trying to get AI past the pilot and into real, owned work, that is the job I do.
Thirteen questions, about three minutes. Answers come straight to me, so the first call starts at the real problem.