1.0.0 shipping now · Apache-2.0

Your Strix Halo box,
running real /v1/*
inference.

Every modality — agents, chat, speech, images, memory — served from a single OpenAI-compatible API on your Strix Halo box or any Linux machine.

install.sh
Linux x86_64 · Python ≥3.12
$curl -fsSL https://hal0.dev/install.sh | bash
Read the docs →★ Star on GitHubApache-2.0 · Linux + systemd
hal0-api:8080 ready·all slots up·GTT 9.2 / 96 GB·probe strix-halo
host strix-halo-01·128 GB UMA · iGPU + XDNA SSE connected
agent
Qwen3-30B-A3B-Instruct-2507-Q4_K_M
serving
142 tok/s
embed
nomic-embed-text-v2-moe-Q4_K_M
serving
116 tok/s
rerank
bge-reranker-v2-m3-q4_k_m
ready
warm
stt
whisper-v3:turbo (NPU)
idle
:8083
tts
kokoro-82M-v1.0
idle
:8084
img
sdxl-turbo
ready
warm
Unified memory (GTT)9.2 / 96 GB carveout · 128 GB DIMM
agent 6.5Gembed 2.3Grerank 0.4Gkernel + ZFS ARC 3.8G
dispatch p50
174ms
/ provider stack

Six engines,
one /v1/* surface.

Chat, embeddings, rerank, speech both ways and image generation — each served by the engine that suits it, on the accelerator that suits it, all behind one OpenAI-compatible API. The picker only offers a backend your hardware can actually honour.

Provider
Workloads
Hardware
Endpoints
llama.cpphal0-combined:0826
chat · embed · rerank · vision
ROCm (default) · Vulkan · CUDA
/v1/chat/v1/embeddings/v1/rerankings
FLMv0.9.44
chat · embed (ASR multiplex)
AMD XDNA NPU (opt-in)
/v1/chat/v1/embeddings
FLM / Whisperwhisper-v3:turbo
speech-to-text (co-loaded with chat)
AMD XDNA NPU
/v1/audio/transcriptions
Moonshinemoonshine-base-en
speech-to-text (cpu device)
CPU
/v1/audio/transcriptions
Kokoro / Qwen3-TTSkokoro-v1 82M · Qwen3-TTS 1.7B
text-to-speech · one-switch CPU⇄GPU swap
CPU (Kokoro) · ROCm (Qwen3-TTS)
/v1/audio/speech
ComfyUIdigest-pinned
image gen · SDXL / SD 1.5 / Flux
ROCm
/v1/images/generations
SLOT LIFECYCLE — ENFORCED IN _transition()
offlinepullingstartingwarmingreadyservingidleunloadingoffline error
/ what's in the box

More than models.
A whole AI stack.

Agents, tools, memory, voice and image generation ship in the box and come up wired to each other — not a model runner you then spend a weekend assembling a platform around.

agents
Agents that live on the box

Hermes installs and bootstraps itself on first run, prewired to the local /v1 API. hal0-brain is the second: a resident operator you ask about the box in plain language, from the dashboard or a headless terminal.

mcp
180 tools over MCP

Two Streamable-HTTP servers mount straight onto the API: models, slots, memory, stacks, bench, telemetry and the updater, all reachable by any MCP client. The 54 destructive ones return pending_approval and wait for a human.

memory
Memory that never leaves the box

Opt-in Hindsight-backed memory with a 26-tool surface. The engine runs embedded and loopback-only, embeddings included, so recall costs nothing and reaches no one. Banks are scoped per agent, and destructive operations write an attributable audit row.

voice
Speech in and out, device-keyed

One stt slot and one tts slot, each resolving by device: Moonshine on CPU or whisper-v3:turbo on the XDNA NPU for listening, Kokoro (CPU) or Qwen3-TTS (GPU) for speaking. Switch the engine, not the plumbing.

images
Image generation on the same API

POST /v1/images/generations runs SDXL, SD 1.5 or Flux through a ComfyUI engine. Generating flips the iGPU into exclusive image mode and back, so chat and images share one accelerator without fighting over it.

prewired
Wired up before you log in

One command brings up the API, a coherent set of slots, Open WebUI on :3001, ComfyUI and the agent — lifecycle-managed as extensions, with cosign-verified updates and one-flag rollback from day one.

/ what people run it for

One box,
three day jobs.

The same install serves all three at once — the slots share the accelerator instead of taking turns with it.

/ local coding stack

Point your editor at your own box

Aim any OpenAI-compatible client at :8080/v1 — chat, completions, and a dedicated coder slot. Your code never leaves the LAN, and there's no per-token bill.

/ private knowledge & RAG

Retrieval grounded on your data

Embeddings, reranking, and a bundled agent with opt-in memory and MCP tools — a full RAG stack that runs on your hardware, not someone else's server.

/ voice & image lab

Speech and images, one control plane

Transcription, text-to-speech, and local image generation behind the same API — STT on the XDNA NPU, TTS switchable between Kokoro (CPU) and Qwen3-TTS (GPU), ComfyUI on the iGPU, switched cleanly so they share the box.

first-party sweep

what this box does

all 26 models · sweep 2026-06-19 →
169.8
fastest decode · tok/s
26
models measured
6248
peak prefill · tok/s
modelparamsdecode ▼prefillgb
qwen3.5-0.8b0.8B169.862480.6
chadrock3-6-35b-uncensored-mtp-strix-lean35B MoE102.189019
chadrock-35b-ace-saber35B-A3B100.590319
qwopus3-5-4b-coder-mtp-q6-k4B85.08893.6
qwen3.6-35b-a3b-crown-halo-mtp-dynamic35B-A3B84.487322.6
258tok/s
primary + embed concurrent
Strix Halo iGPU · ~9 GB GTT
1call
thundering herd, coalesced
single-flight dispatch + cold-cache prefetch
6workloads
chat · embed · rerank · stt · tts · img
as many slots as memory fits
100tok/s
35B model · MTP speculative decode
ace-saber 35B-A3B · 19 GB · single stream
from the roster benchmark →
/ hardware

Strix Halo native.
Not Strix-Halo-only.

The probe is UMA-aware on Strix Halo and falls back to portable parsers on every other host. The dashboard only labels memory "unified" when it actually is. Linux + systemd is the only hard requirement.

reference deployment
Ryzen AI Max+ 395
"Strix Halo" iGPU + XDNA NPU + 128 GB LPDDR5X-8000 unified memory
128GB
unified · BIOS-tunable to ~96 GB GPU
258tok/s
primary + embed concurrent
All published perf numbers come from this box. Q4 70B fits with massive headroom; Q4 MoE 100B+ with a 17–22B active path becomes feasible.
first-class
Ryzen AI Max 385 / 390
Strix Halo with 64 GB unified
64GB
unified memory
~70B
Q4 ceiling, shorter context
Same install path. Every small + mid tier fits; 70B Q4 works with tighter context windows.
experimental
NVIDIA RTX 30 / 40 / 50
10–32 GB dedicated VRAM · CUDA llama.cpp
32GB
RTX 5090 VRAM ceiling
~30B
Q4 comfortable
Same slot lifecycle, dedicated VRAM instead of UMA — higher tok/s on small models, lower ceiling on the big ones.
supported
AMD Radeon RX 7000
16–24 GB discrete · ROCm or Vulkan container profiles
24GB
7900 XTX VRAM
ROCm
· Vulkan
Discrete AMD path, same hal0-slot@<name> lifecycle as Strix Halo — both ROCm and Vulkan container profiles run today.
fallback
CPU-only x86_64
Vulkan-CPU · usable for tiny models
0.5–4B
practical model size
Qwen0.5B
the CI smoke model
CI runs Qwen 0.5B here. Usable for tiny models and smoke tests, not the headline experience.
supported
Proxmox LXC
privileged container · iGPU + XDNA passthrough
0600
PVE token, never echoed back
segmented
host-pressure overlay
Drop a read-only PVEAuditor token into Settings and the memory bar shows physical DIMM total + a muted "Proxmox host" segment for other-tenant + ZFS ARC pressure.
/ vs. the alternatives

The orchestration layer
around your inference engine.

Slots survive hal0-api restarts. Embeddings, rerank, STT, TTS, and image gen all sit behind the same/v1/* surface. UMA-aware hardware probe and slot-fit warnings are first-class, not a slash command in a chat window.

hal0 (this)
ollama
LM Studio
OpenAI cloud
OpenAI-compatible /v1/*
chat · embed · rerank · STT · TTS · img
chat · embed
chat · embed
chat · embed · STT · TTS · img
Concurrent slots
as memory fits
one at a time
one at a time
fully concurrent
Slot lifecycle state machine
typed, atomic, SSE-streamed
no
no
UMA-aware hardware probe
Strix Halo · NPU · platform-aware
CPU / GPU only
CPU / GPU only
XDNA NPU support
first-class via FLM
no
no
Dispatcher with upstreams
OpenRouter · Anthropic · OpenAI · custom
no
no
Cosign-signed updates
keyless OIDC + rollback
no
auto-update
Headless / Linux-first
systemd-required, headless
cross-platform
GUI required
cloud
Your data, your hardware
yes
yes
yes
no
Cost per million tokens
$0 + electricity
$0 + electricity
$0 + electricity
$0.50 – $60

Competitor capabilities reflect each project's published docs as of June 2026 and move fast — treat this as a snapshot, not a live scorecard.

/ the console

Chat, image gen, agents, and memory —
one console.

Slots, models, local image generation, agent memory, an agent task board, MCP, and logs — every surface in one React operator console. SSE-backed, dark by default. Real screenshots from a live hal0 instance.

hal0 dashboard overview — slots, throughput, and live service health
Slots view — per-slot state and the typed inference lifecycle
Slots — one engine, every workload, a typed lifecycle you can watch live.
ComfyUI image generation with the iGPU in exclusive image mode while inference slots are paused
Image gen on the iGPU — generating flips the accelerator into exclusive image mode; inference pauses, then resumes. One GPU, shared cleanly.
Agent memory rendered as a navigable semantic and temporal knowledge graph
Graph memory (opt-in) — facts become a navigable semantic + temporal graph, namespaced per agent.

Plus a Hermes agent that lives on the box, an Operator Board kanban wired to it, an MCP server + client, and a live XDNA NPU view — see theroadmap for everything shipped.

/ agents, tools + memory

Agents that
live on the box.

Hermes installs and bootstraps itself on first run — confined by its own systemd unit, prewired to the local /v1 API and hal0's MCP servers, with opt-in memory. Reach it from Telegram or Discord; it chains tool calls unattended for hours and writes each run back to memory.

  • Self-bootstraps: env probe → model wiring → MCP memory → persona.
  • hal0-agent@hermes.service — NoNewPrivileges, ProtectSystem=strict, a pinned ReadWritePaths set, and the terminal off by default.
  • Gated tools clear an approval bell; every call is audited.
Agents & memory →hover to tilt · click to flip
Hermes — the bundled hal0 agent
Hermes
chadrock-35b-ace-saber
ctx0/164K
remote control · self-improving · orchestration
READYtap · abilities
Hermes · abilities
Ghost Relay40pwr
Reach it from any Telegram or Discord thread.
Engram60pwr
Writes each run back to memory — never relearns it.
Deep Run90pwr
Chains tool calls unattended for hours.
Skills
voice · ttsspeech · sttimage-genvisionembeddings
logspersona
Run agents →
hal0-brain

A resident operator, not a chatbot

hal0's own small always-on model, wired to hal0's admin catalog and reachable from the dashboard's Agent Chat or hal0 chat --brain on a headless box. Ask it what's loaded, what the box is doing, what's eating the memory pool. It ships read-only — it reads and explains out of the box, and one config line lets it act.

147 of the 180 admin tools on its hands · 85 runnable under the shipped default
mcp

The platform is the tool catalog

Two Streamable-HTTP servers mount onto the API —/mcp/admin and /mcp/memory — so any MCP client drives the box: 28 model tools, 26 memory, 22 slots, 17 stacks and profiles, 12 bench, plus telemetry and the updater. A validator enforces zero overlap between gating tiers, so the catalog cannot silently drift.

54 destructive tools queue for a human · 9 that no config can un-gate
memory

Recall that reaches no one

Opt-in memory on a Hindsight engine that runs embedded and loopback-only, local embeddings included — nothing is sent anywhere to remember it. Banks are scoped per agent; recall fans out across the ones a caller can see, then hal0 merges, ranks and fits the result to a token budget.

26 memory tools · destructive ops write an attributable audit row
/ what 1.0 delivers

Everything here runs
on your box today.

The API surface, the slot lifecycle, the security model, install and update, hardware probing, agents and memory — 23 capabilities, every one of them shipped in 1.0. Tagged releases reachreleases.hal0.dev within ~60s.

stable1.0.0cosign-keyless verified · atomic version swap · one-flag rollback

180 admin tools over MCP, slots that report ready only once /health passes,
device-keyed voice — hardened across twelve release candidates.

01
mcp
The whole platform is reachable over MCP
The admin MCP catalog grew from 92 to 180 tools — services, ComfyUI, updater/doctor/health, telemetry, slots, models, bench, approvals — including a 26-tool memory surface at feature parity with Hindsight 0.8.4.
02
slots
Slots start when you say so
A new autoload setting decouples binding a model from boot start, and an eviction priority (0–100) replaces the inert lru opt-in — memory-pressure eviction actually works on a stock box. On-demand capability slots and per-slot profiles round it out.
03
stability
Nothing reports ready until it is
A slot reaches ready only when its own /health passes — never on a systemd snapshot. Transitions are atomic and persisted to state.json, so slots survive an hal0-api restart, and context size is derived per slot instead of silently inheriting the 4096 default.
04
security
Signed, sandboxed, audited
Release tarballs verify against their GitHub OIDC identity before they unpack, and the version swap reverts with one flag. The bundled agent runs under NoNewPrivileges and ProtectSystem=strict with its shell off by default, gated MCP tools clear an approval bell, and every destructive memory op writes an attributable audit row.
Inference + providers
The /v1/* surface and the engines behind it.
6 shipped
OpenAI-compatible /v1/* API
Chat, completions, embeddings, rerank, transcriptions, speech, images. Every OpenAI SDK works unchanged against the local box.
Six-provider stack
llama.cpp (Vulkan / ROCm) for chat and embed, FLM for the XDNA NPU (chat + whisper-v3:turbo STT + embed, one process), Moonshine (CPU) or whisper-v3:turbo (NPU) behind one stt slot, Kokoro (CPU) or Qwen3-TTS (GPU) behind one tts slot, ComfyUI for image generation.
Image generation + iGPU switchover
POST /v1/images/generations served by a ComfyUI engine; generating flips the GPU into exclusive image mode and back, so chat and image share one accelerator.
FLM NPU provider
Self-contained NPU toolbox pinned by digest. The chat + STT + embed trio is surfaced only when XDNA hardware is present.
Bundled chat UI + voice
Open WebUI on :3001 prewired to the local API, whisper-v3:turbo STT (NPU, via FLM), and a one-switch Kokoro (CPU) ⇄ Qwen3-TTS (GPU) TTS engine served through the capability API.
Hands-free voice
Open WebUI's Call mode wired to hal0's /v1/audio/transcriptions and /v1/audio/speech endpoints — whisper-v3:turbo for listen, Kokoro for reply. hal0 also serves its own WS /v1/realtime for clients that want to drive the loop directly.
Stability + slot lifecycle
Atomic transitions, health gated on a real probe, state that survives a restart.
3 shipped
Slot lifecycle state machine
Atomic transitions (offline → pulling → starting → warming → ready → serving), persisted and SSE-streamed.
Honest slot health + derived context
A slot is marked ready only once its real /health passes — never on a systemd snapshot. Context size is derived per slot, never silently inheriting the 4096 default.
Capability slots + profiles
Embed / Voice / Image capability cards and an NPU rollup over flat slots; per-device profiles unify the launch flags, all from one Slots tab.
Install + distribution
One command. Signed. Resumable.
3 shipped
One-command install
install.sh probes the host, picks a hardware-anchored tier, and provisions a coherent set of slots, extensions and models in one pass — resumable, and re-runnable to repair a box in place.
Cosign-signed self-update
Atomic version swap with one-flag rollback. Stable, preview, and nightly channels, GitHub OIDC-verified release tarballs, manifest proxied at releases.hal0.dev.
Extensions framework
Apps and agents (Open WebUI, ComfyUI, Hermes) packaged as auto-wired extensions: lifecycle-managed, selectable at install.
Hardware + observability
Probe before you load. A memory bar that reads the live driver total.
3 shipped
UMA-aware probe
Detects iGPU, XDNA NPU, and the unified memory pool; surfaces fit warnings inline before you load a model that won't fit.
Live GTT total + honest memory bar
The memory bar reports the live GTT total from the driver, not a stale cached probe — so the unified pool you see is the pool you have.
NPU occupancy grid
A living per-slot occupancy view of the XDNA NPU that breathes with real activity instead of a static picker.
Agents + memory + MCP
Agents that run on the box, with memory and tool access.
4 shipped
Bundled Hermes agent
Hermes installs and bootstraps on first run — confined by its own systemd unit, prewired to the local /v1 API and MCP servers, with an agent-card library in the dashboard.
Operator Board
A hal0-skinned kanban wired to Hermes (/api/board/*), with a live agent-chat drawer and working task creation.
Hindsight-backed memory
Opt-in memory engine (HAL0_MEMORY_ENABLED), running embedded and loopback-only. Shared by default; the X-hal0-Agent header scopes a client's writes to its own namespace. Every destructive op — bank delete, memory/document/directive wipes — records a durable, attributable audit row.
MCP host + server
hal0 speaks Model Context Protocol both directions: an admin + memory MCP server, plus an allow-listed client that composes external MCP tools. Destructive calls gate through an approval bell + CLI.
Security
Signed supply chain, contained agents, audited tool calls, LAN-scoped by design.
4 shipped
LAN-scoped by design, proxy at the edge
hal0 ships no auth layer and no TLS of its own. The API binds 0.0.0.0:8080 for the LAN and stops there; the edge is yours — Traefik, nginx, or a Cloudflare Tunnel in front, using the identity provider you already run.
Cosign-keyless supply chain
Every release tarball is verified against its GitHub OIDC identity before it unpacks. The version swap is atomic against a /usr/lib/hal0/current symlink, and --rollback reverts it in one flag.
Contained agent execution
The bundled agent runs under its own systemd unit with NoNewPrivileges, ProtectSystem=strict and a pinned ReadWritePaths set, and its terminal is off by default. Gated tools clear an approval bell before they fire.
Attributable audit trail
Every destructive memory operation — bank delete, memory/document/directive wipes — writes a durable audit row naming the caller, and an x-request-id ties a response back to rows in both the metrics and audit tables.
/ what's next

6 things we intend to build.

No dates and no order. Scope shifts release to release as we learn what people run on a local platform.Open an issue ↗to argue for one.

Fine-tune & LoRA hot-swap
Attach and rotate LoRAs against a warm base model without unloading the underlying weights.
Per-model rate limits & budgets
Cost-style accounting for local inference so a chatty agent can be capped without taking the whole box down.
Benchmarks & presets UI
In-dashboard tok/s + latency runs, plus curated loadout presets you can flash onto a fresh install.
AUR PKGBUILD & Ubuntu PPA
Native distro packages on top of the install script: pacman and apt as first-class install paths.
Multi-host federation
A slot mesh across LAN boxes: chat on the Strix Halo, embed on the workstation, all behind one /v1/* surface.
ChatOps adapters
Slack and Matrix bridges as extensions, so you can talk to hal0 from the rooms you already live in.
forum.hal0.dev

what users are talking about

open the forum ↗
community profiles

optimized model & hardware recipes from the community

all profiles →
learn

latest from the blog

all posts →
/ get hal0

Put the whole stack
on your own box.

One command on a fresh Linux box brings up the API, the slots, voice, image gen, and the agent. Apache-2.0, signed releases, no telemetry.

install.sh
Linux x86_64 · Python ≥3.12
$curl -fsSL https://hal0.dev/install.sh | bash
Apache-2.0Linux + systemdno telemetry by defaultcosign-signed releases