Personal project / AI & systems integration

Sebastian: household AI system

Sebastian is what happened when my home lab stopped being a collection of separate projects and started becoming one system.

I’m connecting the ways we interact with the house to the systems that hold its records. A model saying something happened isn’t enough; I want to know it actually happened. Sebastian is where I learn and experiment with that.

My part

I design the system, define the behavior and boundaries, connect the services, train and evaluate some of the local models, and use AI extensively in the implementation.

Status

Working household system. Some capabilities are deployed, some are experiments, and model/routing work moves through separate evaluation stages before promotion.

The model isn’t the system

Household services hold our records, Home Assistant owns device state, the media stack owns library availability, and the vault holds durable knowledge. Models help interpret requests; deterministic services handle validation, permissions, confirmation, and execution.

InterfacesDiscord · Life Dashboard · Physical interfaces
Conversational / AI layerHermes · Local and cloud models
Deterministic servicesIdentity · Validation · Permissions · Confirmation · Execution
Authoritative records and servicesHousehold services · Home Assistant · Media stack · Vault
A request is not a result.
The media services provide evidence that a movie is actually available to play.
Occupancy is not identity.
A room sensor doesn’t identify who is there. Stale telemetry can’t justify a consequential PC action.
“OK” is not enough.
If an interface reports success but the authoritative record or another interface disagrees, the operation isn’t finished.

What we actually use

My partner and I use the dashboard for tasks, chores, food choices, media, calendar information, and chat. We have separate personal context alongside shared household state, plus a few deliberately silly things like the companion crab*.

*Because who doesn’t miss Tamagotchi, plus he can dance to Crab Rave when we’re having an exceptionally hard day.

The same request, wherever I happen to be

My partner has never seen The Godfather Part II. We will be fixing that.

Sebastian’s watchlist is one small example of how I’ve been building the household system around us. I can use the dashboard when I want to see and manage everything directly, or just tell Sebastian what I want in Discord. Either way, the request ends up in the same shared system.

I also don’t want one request quietly becoming five. Adding a movie to our watchlist doesn’t automatically request it for the media server. Marking it watched doesn’t require a rating. Letterboxd stays a handoff instead of Sebastian posting on my behalf. Those are separate choices because they’re separate actions.

Under the hood

Sebastian is also where I experiment with local models, routing, memory, and agent infrastructure.

Model training & evaluation

Local models and evaluation

I also use Sebastian as a place to train and evaluate local models. The interesting part for me isn’t getting a training run to look good. It’s deciding whether the resulting model behaves well enough to earn a job in the system.

Model evaluation example

One candidate that was not promoted

I trained and evaluated a Qwen3.5 2B LoRA candidate for structured household action proposals. At this evaluation checkpoint, it had improved but had not passed the release requirement.

60 / 60Critical matches on a bounded training-fit panel
72 / 77Subsequent development-test result
76 / 77Frozen requirement for that development test

Broader retention improved from 369/440 to 415/440, but remaining errors included unsafe proposals. The candidate did not qualify for promotion at that checkpoint. The training-fit result did not substitute for the development test or the other release gates.

Earlier evaluation snapshot. These are separate test panels, not a single accuracy measure. Current training and promotion work continues.

Specialist model work continues, with custom datasets, regression testing, and separate evaluation stages. The earlier 3,283-record, 12-task dataset is preparation evidence, not deployment evidence.

Local runtime and training

The model-development work includes Gemma-family adapters, quantization, and cloud training. The October 3, 2026 project review recorded the merged, quantized Sebastian Brain V4 Vision artifact in Ollama. That is dated runtime evidence, not a claim that every model or connector remains in the same state.

Other models serve narrower conversation, compression, perception, OCR, and embedding roles. Some are existing models selected and evaluated for those roles, rather than custom-trained models. On-demand loading and resource limits keep background inference from competing with foreground work or media workloads.

Memory, routing & authority

Memory and improvement

Conversation, durable memory, and operational records have different jobs. Feedback and review queues help identify changes, and repair tooling can produce patches for review. It does not automatically merge or deploy them.

Laya/Jev-style routing is under evaluation. Deterministic software retains authority over permissions, resource limits, confirmation, and execution. A custom Laya router is not deployed.

State and execution

Python domain services keep canonical household records in SQLite. Actions are tied to authenticated users or mapped Discord senders. Stable IDs and revisions address stale clients; idempotency keys prevent retries from repeating an operation. Consequential conversational actions are staged for confirmation where required.

The browser application uses a same-origin API adapter, mobile layouts, and PWA assets. Cross-surface testing includes identity, retries, stale revisions, and household-local day boundaries.

Integrations

Jellyseerr handles media requests, Radarr and Sonarr track fulfillment, and Jellyfin supplies final library-availability evidence. Home Assistant owns device state. Integrations handle unavailable services and preserve manual overrides.

Cardputer, Rabbit R1, and StackChan are part of the physical-interface work. Connector capabilities and end-to-end validation differ across devices.

Everyday interaction

Daily planning limits the work and chores in view, so stale tasks don’t overwhelm the day. Food and media choices narrow through progressive questions, and media reviews ask one question at a time. Everyday views stay separate from administrative and debugging controls.

Local Agent Bridge

Local Agent Bridge

The Local Agent Bridge is a separate Python/FastAPI prototype exploring REST/WebSocket communication, adapters, and per-action confirmation. It is not a complete copy of Sebastian, and external connector delivery remains unfinished.

View the prototype on GitHub