Platform engineering when you can't afford to get it wrong. AI systems that survive production.
Infrastructure written by someone who's shipped it at Alaska Airlines, Expedia, Unity, and — most recently as Senior Staff Architect — in Fortune 500 healthcare.
Kubernetes & cloud-native
Bare-metal to multi-cloud k8s. Production-hardened, Talos-ready, CAPI-provisioned.
Operator & controller dev
Custom operators in Go. controller-runtime, admission webhooks, reconciliation you can trust.
Multi-cloud & hybrid
Active-active across AWS, Azure, GCP — with on-prem, hypervisors, and containers when it matters.
Internal developer platforms
Self-service primitives developers actually use. Backstage, custom portals, whatever the paved road needs.
AI systems & agents
Multi-agent architectures, eval harnesses, local-model pipelines. Built to run offline.
Observability & networking
eBPF-native networking with tracing, metrics, logs, and profiling flowing to one pane of glass.
IaC & GitOps
Infrastructure as real code. Pulumi in Go when the blast radius demands types.
The kind of systems we're actually shipping.
Abstracted from real production architectures. Not slideware. Not AI-generated. Hand-drawn from the runbooks.
harness-go — autonomous multi-agent coding pipeline
Proctor plans. Coders implement in isolated git-worktree generations. Reviewers score and reject. NATS JetStream for orchestration, BadgerDB + Bleve for cross-session memory, hot-reloadable policy via atomic pointer swap during compaction so policy changes land without dropping in-flight work.
ON 128GB STRIX HALO · VULKAN · LOCAL + FRONTIER MODELS
agentisan
Dual-provider sub-agent primitives for Claude Code and Codex, served over MCP and distributed as an OCI artifact alongside a 25+ skill library.
Distributing AI skills as OCI artifacts
What if operational skills — compliance workflows, identity patterns, code-review harnesses — were distributed across a heterogeneous fleet the same way containers are?
Platform Engineering Residency
Senior engineer embedded to build or modernize an IDP, Kubernetes fleet, or developer platform. Heads-down, with ADRs and runbooks as artifacts.
DetailsAI Systems, Built Right
Agent architectures, LLM evaluation harnesses, RAG that works, OCI-distributed skill patterns, multi-agent pipelines. Working systems, not slide decks.
DetailsAI Advisory & Training
Strategy for leaders figuring out where AI actually fits, hands-on workshops, local-model evaluation, and the mechanics of shipping AI in regulated environments.
DetailsInfra Rescue
The “your stack is on fire” short-engagement. Shows up, stabilizes the thing, writes down what to do differently next time.
DetailsOwned hardware means honest claims.
Nerds Run operates its own R&D infrastructure — a 42U rack running Proxmox, Talos, Cilium, and ArgoCD, plus a Strix Halo rig with 128GB unified memory on a Vulkan backend. It's where the multi-agent pipeline, model evaluation harness, and benchmark suite get built and measured before any of it goes near a client.
Not a demo environment. A working shop.
Tell us what you're trying to ship —
we'll tell you what it looks like to get there.
We take one or two new engagements per quarter. No published pricing — every engagement is scope- and duration-dependent, so pricing is a conversation. Bring the problem. We'll bring the questions.