Acquire
Understand CPUs, GPUs, memory, storage, power, thermals, compatibility, and cost.
Source, assemble, boot, benchmark, and optimize a real multi-vendor AI system in a working garage lab. Learn the physical infrastructure behind every model you use.
Train Your Dragon closes that gap. This is not another software tutorial. It is a guided build experience spanning silicon, servers, power, cooling, networking, Linux, runtimes, inference, observability, and economics.
You will move through the same chain of decisions that real infrastructure teams face, from component selection to workload performance.
Understand CPUs, GPUs, memory, storage, power, thermals, compatibility, and cost.
Build a server from components and learn what each connector, rail, fan, and slot actually does.
Configure firmware, install Linux, provision drivers, and bring the hardware to life.
Launch models with production inference software and expose a working endpoint.
Measure latency, throughput, utilization, memory pressure, power, and system behavior.
Find bottlenecks, change the system, and prove whether performance actually improved.
Each module connects the hardware beneath the rack to the software serving the model.
Platform topology, PCIe, memory hierarchy, NUMA, storage, power delivery, airflow, and thermal limits.
Assembly, firmware, BIOS, Linux provisioning, drivers, device discovery, and validation.
Ethernet, topology, bandwidth, latency, local and shared storage, and data movement.
Containers, CUDA and ROCm concepts, PyTorch, model formats, serving engines, and APIs.
Prefill, decode, batching, KV cache, quantization, TTFT, throughput, and tail latency.
Telemetry, utilization, power, failure diagnosis, capacity, cost per token, and optimization tradeoffs.
“The fastest way to understand AI infrastructure is to build it, break it, and make it faster.”
Small cohort. Real equipment. No simulated lab.
You will understand how hardware, software, workload behavior, and economics interact, not as abstract concepts but as a system you personally assembled and operated.
Apply to build yours →Speak confidently about components, constraints, compatibility, and tradeoffs.
Bring up a machine, diagnose failures, and navigate the Linux and GPU stack.
Connect model behavior to compute, memory, networking, power, and cost.
Leave with benchmarks, architecture notes, and a system story you can show.
You do not need to be a hardware expert. You do need curiosity, persistence, and a willingness to get your hands dirty.
Understand what your models and applications demand from the machines beneath them.
Evaluate compute choices, architecture tradeoffs, vendors, and operating costs.
Develop a rare, tangible foundation across systems, hardware, and AI infrastructure.
Required fields are marked with an asterisk.
More cohort details will be shared with selected applicants.
No. The program is designed to make hardware and systems concepts approachable. Basic technical comfort and strong curiosity are more important.
No. Code is part of the experience, but the focus is the full AI compute system, including hardware, Linux, networking, runtimes, performance, power, and cost.
Yes. The defining feature of Train Your Dragon is hands-on work with actual servers, components, accelerators, cables, tools, and software.
The founding cohort is planned for the San Jose and greater Bay Area. Exact dates and venue details will be shared with shortlisted applicants.
No. Cohort size is intentionally limited to preserve the hands-on format. Applications will be reviewed for fit, commitment, and cohort balance.