Table of Contents
| episode | airdate | recorded | status | short_title | full_title | description | summary | guest | guest_title | guest_links | headshot | youtube_url | youtube_id | fireside_player | slug | tags | topics_to_avoid | |||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 020 | 2026-03-31 | published | Zero to Inference | First Boot, First Inference | The IT Guy Show 020 | Damen Knight of CIQ joins me on the hidden cost of configuring Linux for GPU workloads, why general-purpose distros were not built for AI, and what changes. | Damen Knight of CIQ and I dig into the hidden cost of configuring Linux for GPU and AI workloads. | Damen Knight | Sr. Principal Automation Engineer at CIQ | host.png | https://youtu.be/XzW1JmNCYzs | XzW1JmNCYzs | B1g3qpQH+YwgD2E2E | itguyshow-ep20-first-inference |
|
Abstract
In this episode I sit down with Damen Knight, CIQ Sr. Principal Automation Engineer, to talk about something every AI team knows but nobody budgets for: the hidden cost of configuring Linux for GPU workloads.
Together we dig into why general-purpose Linux distributions were never really built for AI, what actually happens when you spend 30 to 60 minutes per node doing manual CUDA setup, and why a "pre-installed" stack and a "validated" stack are two very different things. From driver conflicts and framework dependency failures to the CIQ Linux Kernel and NVIDIA authorization, this episode covers the infrastructure decisions that determine whether your AI program ships fast or gets stuck in configuration hell.
Whether you're an ML engineer tired of fighting CUDA every time a kernel update drops, a sysadmin managing a growing GPU fleet, or a tech leader wondering why production AI is harder than it should be, this one is for you.
Hook
Every AI team pays for it and nobody budgets for it: the time it takes just to make Linux ready for the GPU. Today we get into that hidden cost and how to skip it.
Agenda
Part 1: The State of AI Infrastructure
- Where are most organizations right now with GPU infrastructure?
- What does a typical AI deployment look like before someone like CIQ gets involved?
- How much time do teams spend on infrastructure versus building models?
Part 2: Why General-Purpose Linux Wasn't Built for This
- What makes AI workloads different from a typical server workload?
- Why do mainstream enterprise Linux distributions fall short for GPU work?
- What does manual CUDA setup actually look like, and what goes wrong?
Part 3: What "Validated" Actually Means
- The difference between "pre-installed" and "validated"
- Why PyTorch, CUDA, and DOCA-OFED shipping as a tested unit matters
- The CIQ Linux Kernel and day-one hardware support
Part 4: Live Demo, Bare Metal to First Inference
- Boot RLC Pro AI image, drivers detected automatically
- nvidia-smi and nvcc --version confirm the stack at first boot
- torch.cuda.is_available() confirms GPU access immediately
Keywords
linux, ai infrastructure, cuda, nvidia, gpu, rocky linux, enterprise linux, mlops, devops, machine learning, pytorch, kernel, ciq, rlc pro ai, ai deployment, gpu drivers, hpc, inference
Links
- Linktree: https://linktr.ee/itguyeric
Blog Post
Published to the website: content/blog/itguyshow-ep20-first-inference.md