NERDS RUN LLCWISCONSIN · EST. 2023CRAFTED BY NERDS FOR HUMANS
01SYSTEM ONLINE
R&DOWNED INFRASTRUCTURE
v1.0EST. 2023
SHIPPING PRODUCTION
Nerds Run the World
PLATFORM ENGINEERING · AI SYSTEMS · BUILT TO SURVIVE PRODUCTION

Platform engineering when you can't afford to get it wrong. AI systems that survive production.

Infrastructure written by Alex Banna, who's shipped it at Alaska Airlines, Expedia, Unity, and — most recently as Senior Staff Architect — in Fortune 500 healthcare.

One or two engagements per quarter.SCOPE-DEPENDENT · NO PUBLISHED PRICING
15+
YEARS SHIPPING
10
ENTERPRISE COMPANIES
#4
FLOCI · 23.7K★ OSS
42U
OWNED LAB RACK
128GB
LOCAL INFERENCE
1–2
ENGAGEMENTS / QUARTER
WHERE WE'VE SHIPPEDENTERPRISE · REGULATED · SCALE
GE HealthCareAlaska AirlinesLeverGuaranteed RateNorthwestern MutualUnityExpediaTUNEPoint InsideOptimum EnergyGE HealthCareAlaska AirlinesLeverGuaranteed RateNorthwestern MutualUnityExpediaTUNEPoint InsideOptimum Energy
03 · SYSTEM DESIGN · LIVE

The kind of systems we're actually shipping.

Abstracted from real production architectures. Not slideware. Not AI-generated. Hand-drawn from the runbooks.

HA Kubernetes architectureA highly available Kubernetes cluster. Client DNS resolves through Route 53 latency routing to a cross-zone network load balancer. The control plane spans three availability zones: three kube-apiserver replicas, each paired with one member of a stacked three-node etcd cluster running Raft, so quorum is two and the cluster survives the loss of one zone. Scheduler and controller-manager run with leader election; the cloud controller manager runs the node, service load balancer and route controllers. The data plane is four Talos Linux worker nodes running Cilium, reached by XDP and ECMP, each carrying a set of pods. A platform column holds ArgoCD for GitOps, an OpenTelemetry collector feeding Prometheus, Loki, Tempo, Mimir and Grafana, and Vault, Kyverno and External Secrets for policy and secret management. ArgoCD syncs to the API servers. // HA KUBERNETES · STACKED ETCD · 3 CONTROL / 4 WORKERS CLIENT / DNS route53 · latency EXTERNAL LB NLB · cross-zone CONTROL PLANE 3× AZ · quorum=2 kube-apiserver AZ-a kube-apiserver AZ-b kube-apiserver AZ-c etcd-0 raft · follower etcd-1 raft · leader etcd-2 raft · follower raft consensus scheduler + controller-mgr leader-election cloud-ccm node · LB · route ctrl DATA PLANE cilium eBPF · 4 nodes node-0 talos · cilium pods node-1 talos · cilium pods node-2 talos · cilium pods node-3 talos · cilium pods ↓ XDP · ECMP PLATFORM GitOps · Obs ArgoCD gitops · app-of-apps OpenTelemetry collector · gateway Prom metrics Loki logs Tempo traces Mimir long-term Grafana unified observability Vault CSI Kyverno policy External-Secrets sync · rotate sync * request hits nearest PoP← quorum survives 1 AZ lossTALOS LINUX · CILIUM/HUBBLE · CLUSTER API · ARGOCD · OTEL · KYVERNO · VAULT · EXTERNAL-SECRETS
Control planeData plane / liveState / data flowPolicy / governance· click a tab to switch scene

Owned hardware means honest claims.

Nerds Run operates its own R&D infrastructure — a 42U rack running Proxmox, Talos, Cilium, and ArgoCD, plus a Strix Halo rig with 128GB unified memory on a Vulkan backend. It's where the multi-agent pipeline, model evaluation harness, and benchmark suite get built and measured before any of it goes near a client.

Not a demo environment. A working shop.

Compute
AMD Strix Halo · 128 GB
Rack
42U · Proxmox · Talos
Network
Cilium · eBPF · 10 GbE
Runtime
Vulkan · llama.cpp · vLLM
rack · nerds run hq · 0142U · live
Let's talk

Tell us what you're trying to ship
we'll tell you what it looks like to get there.

We take one or two new engagements per quarter. No published pricing — every engagement is scope- and duration-dependent, so pricing is a conversation. Bring the problem. We'll bring the questions.