Personal project / AI & systems integration
Sebastian: household AI system
Sebastian is what happened when my home lab stopped being a collection of separate projects and started becoming one system.
I’m connecting the ways we interact with the house to the systems that hold its records. A model saying something happened isn’t enough; I want to know it actually happened. Sebastian is where I learn and experiment with that.
The model isn’t the system
Household services hold our records, Home Assistant owns device state, the media stack owns library availability, and the vault holds durable knowledge. Models help interpret requests; deterministic services handle validation, permissions, confirmation, and execution.
- A request is not a result.
- The media services provide evidence that a movie is actually available to play.
- Occupancy is not identity.
- A room sensor doesn’t identify who is there. Stale telemetry can’t justify a consequential PC action.
- “OK” is not enough.
- If an interface reports success but the authoritative record or another interface disagrees, the operation isn’t finished.
What we actually use
My partner and I use the dashboard for tasks, chores, food choices, media, calendar information, and chat. We have separate personal context alongside shared household state, plus a few deliberately silly things like the companion crab*.
*Because who doesn’t miss Tamagotchi, plus he can dance to Crab Rave when we’re having an exceptionally hard day.
The same request, wherever I happen to be
My partner has never seen The Godfather Part II. We will be fixing that.
Sebastian’s watchlist is one small example of how I’ve been building the household system around us. I can use the dashboard when I want to see and manage everything directly, or just tell Sebastian what I want in Discord. Either way, the request ends up in the same shared system.
I also don’t want one request quietly becoming five. Adding a movie to our watchlist doesn’t automatically request it for the media server. Marking it watched doesn’t require a rating. Letterboxd stays a handoff instead of Sebastian posting on my behalf. Those are separate choices because they’re separate actions.
Under the hood
Sebastian is also where I experiment with local models, routing, memory, and agent infrastructure.
Model training & evaluation
Local models and evaluation
I also use Sebastian as a place to train and evaluate local models. The interesting part for me isn’t getting a training run to look good. It’s deciding whether the resulting model behaves well enough to earn a job in the system.
Model evaluation example
One candidate that was not promoted
I trained and evaluated a Qwen3.5 2B LoRA candidate for structured household action proposals. At this evaluation checkpoint, it had improved but had not passed the release requirement.
Broader retention improved from 369/440 to 415/440, but remaining errors included unsafe proposals. The candidate did not qualify for promotion at that checkpoint. The training-fit result did not substitute for the development test or the other release gates.
Earlier evaluation snapshot. These are separate test panels, not a single accuracy measure. Current training and promotion work continues.
Specialist model work continues, with custom datasets, regression testing, and separate evaluation stages. The earlier 3,283-record, 12-task dataset is preparation evidence, not deployment evidence.
Local runtime and training
The model-development work includes Gemma-family adapters, quantization, and cloud training. The October 3, 2026 project review recorded the merged, quantized Sebastian Brain V4 Vision artifact in Ollama. That is dated runtime evidence, not a claim that every model or connector remains in the same state.
Other models serve narrower conversation, compression, perception, OCR, and embedding roles. Some are existing models selected and evaluated for those roles, rather than custom-trained models. On-demand loading and resource limits keep background inference from competing with foreground work or media workloads.
Memory, routing & authority
Memory and improvement
Conversation, durable memory, and operational records have different jobs. Feedback and review queues help identify changes, and repair tooling can produce patches for review. It does not automatically merge or deploy them.
Laya/Jev-style routing is under evaluation. Deterministic software retains authority over permissions, resource limits, confirmation, and execution. A custom Laya router is not deployed.
State and execution
Python domain services keep canonical household records in SQLite. Actions are tied to authenticated users or mapped Discord senders. Stable IDs and revisions address stale clients; idempotency keys prevent retries from repeating an operation. Consequential conversational actions are staged for confirmation where required.
The browser application uses a same-origin API adapter, mobile layouts, and PWA assets. Cross-surface testing includes identity, retries, stale revisions, and household-local day boundaries.
Integrations
Jellyseerr handles media requests, Radarr and Sonarr track fulfillment, and Jellyfin supplies final library-availability evidence. Home Assistant owns device state. Integrations handle unavailable services and preserve manual overrides.
Cardputer, Rabbit R1, and StackChan are part of the physical-interface work. Connector capabilities and end-to-end validation differ across devices.
Everyday interaction
Daily planning limits the work and chores in view, so stale tasks don’t overwhelm the day. Food and media choices narrow through progressive questions, and media reviews ask one question at a time. Everyday views stay separate from administrative and debugging controls.
Local Agent Bridge
Local Agent Bridge
The Local Agent Bridge is a separate Python/FastAPI prototype exploring REST/WebSocket communication, adapters, and per-action confirmation. It is not a complete copy of Sebastian, and external connector delivery remains unfinished.
View the prototype on GitHub