$ cat projects/local-ai-platform.md

Local-first AI platform

Model serving, memory and agent governance that keep working with the network cable out

in production Aug 2024 — Present Architect and operator

at a glance

Constraint
Offline-capable by design
Routing
By task complexity and cost, not by model size
Memory
Bi-temporal — 750k+ observations under validity versioning
Governance
75+ skill modules · 100+ deterministic hooks
Benchmarking
Own harness, pre-registered bars

context

A platform built for my own work and used daily: model serving, routing, memory and agent execution, all self-hosted.

the problem

Hosted inference means per-token cost, data leaving the machine, and a hard dependency on someone else's uptime. For continuous agent workloads that is three separate problems wearing one coat.

The goal was narrow: a stack that keeps working with the network cable out, and that can say precisely why it produced a particular answer.

approach

What was tried — and, where it applies, what it taught. The second half is the part that usually gets edited out, and the part that is actually useful.

    • A unified gateway in front of multiple backends, routing by task complexity and cost.

    learned Routing by model size was the first attempt and it was wrong. The cheapest adequate path usually beats the strongest available one, and proving which is which required building a benchmark harness before the router could be trusted.

    • Bi-temporal memory: facts carry valid-from and valid-to, so the store can answer "what was true then" and not only "what is true now" — over 750,000 observations handled this way.

    learned A single timestamp answers the wrong question. When a decision is audited months later, what matters is what the system believed at the time — which a mutable record quietly destroys.

    • Agent governance as deterministic rules evaluated before actions — 75+ standardised, model-agnostic skill modules and 100+ executable hooks — rather than instructions inside a prompt.

    learned Anything enforced only by prompt text is a suggestion. The rules that matter had to move out of the model and into code that can refuse.

    • Measured the whole thing — capacity, latency, cost and energy — instead of trusting benchmarks from publishers.

    learned Published benchmarks measure a different workload on different hardware. The only numbers that decide anything are the ones measured on the path actually used.

outcome

Availability
Fully offline-capable
no required external dependency at runtime
Evaluation
Own harness
12 domains, fixed task list, pre-registered bars
Energy
Tracked as a first-class dimension
tokens per joule, measured per configuration
Governance
Deterministic rules
structured modules plus executable guards

what I would do differently

Local-first is a set of constraints, not a virtue. It costs real engineering time and buys independence — worth it for a continuous workload, questionable for a weekly batch job.

The other honest finding: the benchmark became more valuable than any individual model it evaluated. Capacity knowledge that has an evidence pointer behind it can refuse to run a configuration; a leaderboard cannot.

artifacts

No public artifact: this work lives in private repositories. The description above is the citable summary.