<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:podcast="https://podcastindex.org/namespace/1.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
<channel>
<title>NVIDIA Generative AI</title>
<description>What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists.

Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave.

NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did.

This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people.

If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.</description>
<itunes:summary>What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists.

Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave.

NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did.

This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people.

If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.</itunes:summary>
<itunes:subtitle>From language models to cameras, cars, and robots</itunes:subtitle>
<itunes:keywords>generative AI, NVIDIA, large language models, transformers, diffusion models, multimodal AI, fine-tuning, retrieval augmented generation, GPU computing, physical AI</itunes:keywords>
<link>https://www.cloudadorn.com/</link>
<language>en-us</language>
<copyright>Cloudadorn Academy</copyright>
<atom:link href="https://media.cloudadorn.org/nvidia-generative-ai/feed.xml" rel="self" type="application/rss+xml"/>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:category text="Technology"/>
<itunes:category text="Education">
<itunes:category text="Courses"/>
</itunes:category>
<itunes:explicit>false</itunes:explicit>
<itunes:author>Cloudadorn Academy</itunes:author>
<itunes:owner>
<itunes:name>Cloudadorn Academy</itunes:name>
<itunes:email>info@cloudadorn.com</itunes:email>
</itunes:owner>
<itunes:type>serial</itunes:type>
<itunes:complete>yes</itunes:complete>
<podcast:guid>0e702681-1868-5d02-a410-9e1047054e05</podcast:guid>
<podcast:locked owner="info@cloudadorn.com">no</podcast:locked>
<lastBuildDate>Tue, 25 Aug 2026 18:00:00 +0000</lastBuildDate>
<content:encoded><![CDATA[<p>What is actually inside the phrase generative AI, and why does every answer end up being a cost decision? NVIDIA Generative AI is a twenty-three-episode series from Cloudadorn Academy for people who want the real picture and are not already specialists.</p><p>Shelley asks the questions a curious adult would actually ask. Rob answers with something you can picture first and the name of the idea second: nesting dolls for the AI hierarchy, a group chat for attention, sand and statues for diffusion, a case conference for sensor fusion. Season one walks four acts. What these things are, from the hierarchy and the learning paradigms through attention, training, and the architecture zoo. How you make a model yours for the least money that works, from prompting to retrieval to LoRA to a full fine-tune, and which number tells you it worked. How a model gets eyes and ears, through shared embedding spaces, patches as tokens, diffusion, segmentation, and sensor fusion. Then the stack itself: precision formats and memory bandwidth, serving and packaging, the training pipeline, agents, guardrails, robots and world models, and why the ecosystem is hard to leave.</p><p>NVIDIA is the through-line because it sells at every one of those layers, so its product map doubles as a map of the field. Where a vendor number cannot be confirmed against a primary source, the hosts hedge it or drop it and say which one they did.</p><p>This series is narrated using AI voice technology. The content and scripts are original. Shelley and Rob are original hosts, not impersonations of real people.</p><p>If the field finally holds still long enough to make sense, subscribe, leave a review where you listen, and visit cloudadorn.com.</p>]]></content:encoded>
<item>
<title>S1E01. Nesting Dolls: What Generative AI Actually Contains</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e01</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-01.mp3" length="26289248" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:00:00 +0000</pubDate>
<itunes:duration>1057</itunes:duration>
<itunes:episode>1</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>generative AI, machine learning, deep learning, foundation models, large language models, self-supervised learning</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-01.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Artificial intelligence, machine learning, deep learning, generative AI. Four phrases, used interchangeably, by the same person, about the same product. They are not synonyms. They are nested, and every step inward is a stronger claim about what is actually happening.

Shelley makes Rob open the dolls one at a time, biggest to smallest. You will learn why the nesting only runs one way, so a chess program built from hand-written rules is AI and is not machine learning. Why foundation model is not a fifth doll at all: it describes a model trained wide enough to be reused, which is not the same as a model that makes things, and CLIP is the counterexample that proves it — broad, reusable, and it makes nothing at all. NVIDIA&apos;s Cosmos family is the other half of the point: world foundation models that generate video of the physical world, also not language models. The difference between learning the boundary between categories and learning the shape of the data, which is the whole reason a model can make something new instead of only sorting what exists. And the four ways a model learns, including the one people get backwards: pretraining a large language model is self-supervised, not unsupervised, because hiding the next word turns the sentence into its own answer key.

Then the arc. 2012, when a deep network won the ImageNet competition by a distance. 2017 and the paper that threw out recurrence. 2020 and a model doing new tasks from examples in the prompt. 2021, when the same architecture came for pictures. Underneath all of it, multiplying big grids of numbers on a chip that was built for video games.

Takeaway: when a product says AI, it has told you almost nothing. Ask which doll.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Artificial intelligence, machine learning, deep learning, generative AI. Four phrases, used interchangeably, by the same person, about the same product. They are not synonyms. They are nested, and every step inward is a stronger claim about what is actually happening.

Shelley makes Rob open the dolls one at a time, biggest to smallest. You will learn why the nesting only runs one way, so a chess program built from hand-written rules is AI and is not machine learning. Why foundation model is not a fifth doll at all: it describes a model trained wide enough to be reused, which is not the same as a model that makes things, and CLIP is the counterexample that proves it — broad, reusable, and it makes nothing at all. NVIDIA&apos;s Cosmos family is the other half of the point: world foundation models that generate video of the physical world, also not language models. The difference between learning the boundary between categories and learning the shape of the data, which is the whole reason a model can make something new instead of only sorting what exists. And the four ways a model learns, including the one people get backwards: pretraining a large language model is self-supervised, not unsupervised, because hiding the next word turns the sentence into its own answer key.

Then the arc. 2012, when a deep network won the ImageNet competition by a distance. 2017 and the paper that threw out recurrence. 2020 and a model doing new tasks from examples in the prompt. 2021, when the same architecture came for pictures. Underneath all of it, multiplying big grids of numbers on a chip that was built for video games.

Takeaway: when a product says AI, it has told you almost nothing. Ask which doll.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Artificial intelligence, machine learning, deep learning, generative AI. Four phrases, used interchangeably, by the same person, about the same product. They are not synonyms. They are nested, and every step inward is a stronger claim about what is actually happening.</p><p>Shelley makes Rob open the dolls one at a time, biggest to smallest. You will learn why the nesting only runs one way, so a chess program built from hand-written rules is AI and is not machine learning. Why foundation model is not a fifth doll at all: it describes a model trained wide enough to be reused, which is not the same as a model that makes things, and CLIP is the counterexample that proves it — broad, reusable, and it makes nothing at all. NVIDIA's Cosmos family is the other half of the point: world foundation models that generate video of the physical world, also not language models. The difference between learning the boundary between categories and learning the shape of the data, which is the whole reason a model can make something new instead of only sorting what exists. And the four ways a model learns, including the one people get backwards: pretraining a large language model is self-supervised, not unsupervised, because hiding the next word turns the sentence into its own answer key.</p><p>Then the arc. 2012, when a deep network won the ImageNet competition by a distance. 2017 and the paper that threw out recurrence. 2020 and a model doing new tasks from examples in the prompt. 2021, when the same architecture came for pictures. Underneath all of it, multiplying big grids of numbers on a chip that was built for video games.</p><p>Takeaway: when a product says AI, it has told you almost nothing. Ask which doll.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E02. Attention: How a Model Reads a Sentence</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e02</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-02.mp3" length="25498271" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:02:00 +0000</pubDate>
<itunes:duration>1024</itunes:duration>
<itunes:episode>2</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>transformer, attention mechanism, query key value, mixture of experts, encoder decoder, long context</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-02.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

You type a sentence. Something types back. This is what happens to your words in between, in order, with the clever bit slowed right down.

Rob walks Shelley through the five steps every one of these models runs: the sentence is chopped into pieces smaller than words, each piece becomes a location in a space, position gets added on purpose, the stack of blocks runs, and one word comes out. Then it all runs again. The step that sounds like housekeeping turns out to be the one worth stopping on: attention is blind to order. It sees a heap, not a line. Leave that step out and &quot;the dog bit the man&quot; and &quot;the man bit the dog&quot; are the same input to the model.

Then attention itself, which is smaller than its reputation. Every token asks a question, advertises what it knows, and offers something to contribute. Everything gets weighed against everything else, and nothing is ever ignored, only weighted near zero. That last part comes with a bill: ten pieces is a hundred comparisons, twenty is four hundred. Double the length, quadruple the work. That single curve is why long context is expensive, and every trick in the back half of this episode is somebody trying to get out from under it.

Also: why a blindfold is the reason text models are built the way they are, how to tell a model&apos;s shape from the job it does, and why one model honestly has two parameter counts ten times apart. NVIDIA&apos;s Nemotron 3 Nano has 31.6 billion parameters and runs 3.2 billion of them per token. When somebody quotes you a parameter count, ask which one.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

You type a sentence. Something types back. This is what happens to your words in between, in order, with the clever bit slowed right down.

Rob walks Shelley through the five steps every one of these models runs: the sentence is chopped into pieces smaller than words, each piece becomes a location in a space, position gets added on purpose, the stack of blocks runs, and one word comes out. Then it all runs again. The step that sounds like housekeeping turns out to be the one worth stopping on: attention is blind to order. It sees a heap, not a line. Leave that step out and &quot;the dog bit the man&quot; and &quot;the man bit the dog&quot; are the same input to the model.

Then attention itself, which is smaller than its reputation. Every token asks a question, advertises what it knows, and offers something to contribute. Everything gets weighed against everything else, and nothing is ever ignored, only weighted near zero. That last part comes with a bill: ten pieces is a hundred comparisons, twenty is four hundred. Double the length, quadruple the work. That single curve is why long context is expensive, and every trick in the back half of this episode is somebody trying to get out from under it.

Also: why a blindfold is the reason text models are built the way they are, how to tell a model&apos;s shape from the job it does, and why one model honestly has two parameter counts ten times apart. NVIDIA&apos;s Nemotron 3 Nano has 31.6 billion parameters and runs 3.2 billion of them per token. When somebody quotes you a parameter count, ask which one.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>You type a sentence. Something types back. This is what happens to your words in between, in order, with the clever bit slowed right down.</p><p>Rob walks Shelley through the five steps every one of these models runs: the sentence is chopped into pieces smaller than words, each piece becomes a location in a space, position gets added on purpose, the stack of blocks runs, and one word comes out. Then it all runs again. The step that sounds like housekeeping turns out to be the one worth stopping on: attention is blind to order. It sees a heap, not a line. Leave that step out and "the dog bit the man" and "the man bit the dog" are the same input to the model.</p><p>Then attention itself, which is smaller than its reputation. Every token asks a question, advertises what it knows, and offers something to contribute. Everything gets weighed against everything else, and nothing is ever ignored, only weighted near zero. That last part comes with a bill: ten pieces is a hundred comparisons, twenty is four hundred. Double the length, quadruple the work. That single curve is why long context is expensive, and every trick in the back half of this episode is somebody trying to get out from under it.</p><p>Also: why a blindfold is the reason text models are built the way they are, how to tell a model's shape from the job it does, and why one model honestly has two parameter counts ten times apart. NVIDIA's Nemotron 3 Nano has 31.6 billion parameters and runs 3.2 billion of them per token. When somebody quotes you a parameter count, ask which one.</p>]]></content:encoded>
</item>
<item>
<title>S1E03. How a Model Learns: Forward Pass, Loss, and Backpropagation</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e03</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-03.mp3" length="25455908" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:04:00 +0000</pubDate>
<itunes:duration>1022</itunes:duration>
<itunes:episode>3</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>backpropagation, loss function, learning rate, gradient descent, activation function, hyperparameters</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-03.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Training a model is four steps in a loop: guess, score the guess, work out who is to blame, nudge everybody, then do it again. When you type into a finished model and it answers, that is the forward pass, and the forward pass changes nothing. Learning is the backward pass plus an optimizer step. A model out in the world is frozen.

Shelley makes Rob slow the blame step down. One number says how wrong the answer was. Billions of numbers inside contributed to it. Backpropagation works out, for every one of them, how much of that error was its fault, by going backwards a layer at a time. The loss you pick is you telling the machine what you care about: next-word prediction is classification over the whole vocabulary. The crude activation beat the sophisticated ones because it does not destroy the correction on the way back. Residual connections are a megaphone so the message is not whispered through sixty handovers. Six knobs, each of which breaks training in its own way, and the distinction that is the spine of the last third: parameters are learned. Hyperparameters are chosen.

Takeaway: nothing is learned on the way forwards.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Training a model is four steps in a loop: guess, score the guess, work out who is to blame, nudge everybody, then do it again. When you type into a finished model and it answers, that is the forward pass, and the forward pass changes nothing. Learning is the backward pass plus an optimizer step. A model out in the world is frozen.

Shelley makes Rob slow the blame step down. One number says how wrong the answer was. Billions of numbers inside contributed to it. Backpropagation works out, for every one of them, how much of that error was its fault, by going backwards a layer at a time. The loss you pick is you telling the machine what you care about: next-word prediction is classification over the whole vocabulary. The crude activation beat the sophisticated ones because it does not destroy the correction on the way back. Residual connections are a megaphone so the message is not whispered through sixty handovers. Six knobs, each of which breaks training in its own way, and the distinction that is the spine of the last third: parameters are learned. Hyperparameters are chosen.

Takeaway: nothing is learned on the way forwards.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Training a model is four steps in a loop: guess, score the guess, work out who is to blame, nudge everybody, then do it again. When you type into a finished model and it answers, that is the forward pass, and the forward pass changes nothing. Learning is the backward pass plus an optimizer step. A model out in the world is frozen.</p><p>Shelley makes Rob slow the blame step down. One number says how wrong the answer was. Billions of numbers inside contributed to it. Backpropagation works out, for every one of them, how much of that error was its fault, by going backwards a layer at a time. The loss you pick is you telling the machine what you care about: next-word prediction is classification over the whole vocabulary. The crude activation beat the sophisticated ones because it does not destroy the correction on the way back. Residual connections are a megaphone so the message is not whispered through sixty handovers. Six knobs, each of which breaks training in its own way, and the distinction that is the spine of the last third: parameters are learned. Hyperparameters are chosen.</p><p>Takeaway: nothing is learned on the way forwards.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E04. When Training Goes Wrong: Overfitting, Bad Data, and Metrics That Lie</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e04</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-04.mp3" length="25175813" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:06:00 +0000</pubDate>
<itunes:duration>1011</itunes:duration>
<itunes:episode>4</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>overfitting, underfitting, class imbalance, precision and recall, F1 score, exploratory data analysis</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-04.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start.

Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone.

Takeaway: the most expensive mistake is treating the wrong failure.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start.

Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone.

Takeaway: the most expensive mistake is treating the wrong failure.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>There are two ways for training to fail, they are opposites, and the fix for one makes the other worse. Overfitting is the driver who learned one route to work perfectly and is helpless on any other street. Underfitting is the driver who had one lesson and stopped. The only way to tell which you have is a slice of data locked in a drawer before you start.</p><p>Rob walks Shelley through the pair of numbers that is the entire diagnosis, early stopping done honestly on a bumpy curve, and why looking at the data comes first. Label errors in the collections this industry measures itself against. Five data failures, each with its own symptom: lopsided categories, gaps, extremes, leakage, duplicates. Accuracy is the number that lies most often when the rare thing is the whole job. Precision and recall point at two different mistakes, and each can be gamed alone.</p><p>Takeaway: the most expensive mistake is treating the wrong failure.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E05. The Architecture Zoo: CNN, RNN, Transformer, GAN, VAE, Diffusion</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e05</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-05.mp3" length="22826846" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:08:00 +0000</pubDate>
<itunes:duration>913</itunes:duration>
<itunes:episode>5</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>neural network architectures, CNN, GAN, VAE, diffusion models, latent space</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-05.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

These are not brands. They are shapes. Somebody looked hard at one kind of problem, worked out what shape of machine would suit it, and the shape got a name. Every name on the shelf is superb at one thing sitting directly next to a thing it is hopeless at.

Rob names them that way, and no deeper: the CNN sliding a window over a grid, the RNN walking a sequence and forgetting, LSTM holding on longer, the transformer looking at every position at once and paying quadratic rent for it, Mamba as the young challenger whose work grows with length instead of length squared. Then the makers: a GAN is two networks set against each other, and mode collapse is the word that belongs to that family only. The autoencoder squeezes a thing down to a recipe. The VAE lets you walk around inside that recipe. Diffusion gets named and handed forward. Latent space is the idea worth taking with you.

Takeaway: ask what the shape is for, and what it cannot do.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

These are not brands. They are shapes. Somebody looked hard at one kind of problem, worked out what shape of machine would suit it, and the shape got a name. Every name on the shelf is superb at one thing sitting directly next to a thing it is hopeless at.

Rob names them that way, and no deeper: the CNN sliding a window over a grid, the RNN walking a sequence and forgetting, LSTM holding on longer, the transformer looking at every position at once and paying quadratic rent for it, Mamba as the young challenger whose work grows with length instead of length squared. Then the makers: a GAN is two networks set against each other, and mode collapse is the word that belongs to that family only. The autoencoder squeezes a thing down to a recipe. The VAE lets you walk around inside that recipe. Diffusion gets named and handed forward. Latent space is the idea worth taking with you.

Takeaway: ask what the shape is for, and what it cannot do.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>These are not brands. They are shapes. Somebody looked hard at one kind of problem, worked out what shape of machine would suit it, and the shape got a name. Every name on the shelf is superb at one thing sitting directly next to a thing it is hopeless at.</p><p>Rob names them that way, and no deeper: the CNN sliding a window over a grid, the RNN walking a sequence and forgetting, LSTM holding on longer, the transformer looking at every position at once and paying quadratic rent for it, Mamba as the young challenger whose work grows with length instead of length squared. Then the makers: a GAN is two networks set against each other, and mode collapse is the word that belongs to that family only. The autoencoder squeezes a thing down to a recipe. The VAE lets you walk around inside that recipe. Diffusion gets named and handed forward. Latent space is the idea worth taking with you.</p><p>Takeaway: ask what the shape is for, and what it cannot do.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E06. The Cheapest Thing That Works: Prompting Before Fine-Tuning</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e06</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-06.mp3" length="24310236" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:10:00 +0000</pubDate>
<itunes:duration>975</itunes:duration>
<itunes:episode>6</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>prompt engineering, few-shot prompting, chain of thought, model customization, temperature top-k top-p, ReAct</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-06.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

The cheapest thing that works is a sentence, not a training run. Before you fine-tune, you try talking to the model you already have. Prompting is that: you change what you say, not the weights.

Shelley makes Rob put the ladder in order. Zero-shot, one example, a handful of examples. Chain of thought, which is asking it to show its working. ReAct, which is think then look something up then think again, and which must not smash the English verb react. Temperature, top-k, top-p: three knobs for how adventurous the next word is allowed to be. A prompt is not a program. It is not reliable the way a fine-tune can be. The honest use of this hour is knowing when the cheap move is enough and when it is theatre.

Takeaway: try the sentence before you pay for the training run.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

The cheapest thing that works is a sentence, not a training run. Before you fine-tune, you try talking to the model you already have. Prompting is that: you change what you say, not the weights.

Shelley makes Rob put the ladder in order. Zero-shot, one example, a handful of examples. Chain of thought, which is asking it to show its working. ReAct, which is think then look something up then think again, and which must not smash the English verb react. Temperature, top-k, top-p: three knobs for how adventurous the next word is allowed to be. A prompt is not a program. It is not reliable the way a fine-tune can be. The honest use of this hour is knowing when the cheap move is enough and when it is theatre.

Takeaway: try the sentence before you pay for the training run.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>The cheapest thing that works is a sentence, not a training run. Before you fine-tune, you try talking to the model you already have. Prompting is that: you change what you say, not the weights.</p><p>Shelley makes Rob put the ladder in order. Zero-shot, one example, a handful of examples. Chain of thought, which is asking it to show its working. ReAct, which is think then look something up then think again, and which must not smash the English verb react. Temperature, top-k, top-p: three knobs for how adventurous the next word is allowed to be. A prompt is not a program. It is not reliable the way a fine-tune can be. The honest use of this hour is knowing when the cheap move is enough and when it is theatre.</p><p>Takeaway: try the sentence before you pay for the training run.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E07. Teaching an Old Model New Tricks: Transfer Learning, LoRA, and RLHF</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e07</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-07.mp3" length="25147738" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:12:00 +0000</pubDate>
<itunes:duration>1009</itunes:duration>
<itunes:episode>7</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>transfer learning, fine-tuning, LoRA, QLoRA, RLHF, knowledge distillation</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-07.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

You do not build one of these from scratch, and the reason is not only money. You take a model that already exists and reuse what it already learned. The early layers are general. The specific part sits at the far end, near the answer. That reuse is transfer learning.

Rob sorts five stages by how much of your own material they need and who in the world actually runs them. Original training on a web-scale pile, next word over and over, self-supervised. Continued pretraining on one field. Supervised fine-tuning on written pairs, which is the first stage a normal organisation actually runs. Then preference alignment: you cannot write down the right answer to kindly, but you can pick the better of two attempts in a second. That loop is RLHF. Finishing school. A much smaller model put through it was preferred to a far bigger one that was not.

Then the cheap adapters. PEFT. LoRA trains a small piece and leaves the rest frozen. QLoRA squashes the frozen copy first. Distillation teaches a small model to imitate a large one.

Takeaway: somebody else paid for the years; you pay for the menu.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

You do not build one of these from scratch, and the reason is not only money. You take a model that already exists and reuse what it already learned. The early layers are general. The specific part sits at the far end, near the answer. That reuse is transfer learning.

Rob sorts five stages by how much of your own material they need and who in the world actually runs them. Original training on a web-scale pile, next word over and over, self-supervised. Continued pretraining on one field. Supervised fine-tuning on written pairs, which is the first stage a normal organisation actually runs. Then preference alignment: you cannot write down the right answer to kindly, but you can pick the better of two attempts in a second. That loop is RLHF. Finishing school. A much smaller model put through it was preferred to a far bigger one that was not.

Then the cheap adapters. PEFT. LoRA trains a small piece and leaves the rest frozen. QLoRA squashes the frozen copy first. Distillation teaches a small model to imitate a large one.

Takeaway: somebody else paid for the years; you pay for the menu.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>You do not build one of these from scratch, and the reason is not only money. You take a model that already exists and reuse what it already learned. The early layers are general. The specific part sits at the far end, near the answer. That reuse is transfer learning.</p><p>Rob sorts five stages by how much of your own material they need and who in the world actually runs them. Original training on a web-scale pile, next word over and over, self-supervised. Continued pretraining on one field. Supervised fine-tuning on written pairs, which is the first stage a normal organisation actually runs. Then preference alignment: you cannot write down the right answer to kindly, but you can pick the better of two attempts in a second. That loop is RLHF. Finishing school. A much smaller model put through it was preferred to a far bigger one that was not.</p><p>Then the cheap adapters. PEFT. LoRA trains a small piece and leaves the rest frozen. QLoRA squashes the frozen copy first. Distillation teaches a small model to imitate a large one.</p><p>Takeaway: somebody else paid for the years; you pay for the menu.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E08. Open Book: RAG Done Honestly</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e08</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-08.mp3" length="24889537" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:14:00 +0000</pubDate>
<itunes:duration>999</itunes:duration>
<itunes:episode>8</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>retrieval augmented generation, RAG, embeddings, vector database, chunking, NeMo Retriever</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-08.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Open book. The model does not have to remember your documents. It has to find the right page and read it at the moment you ask. That is retrieval augmented generation, and the honest version is smaller than the demo.

Rob walks the pipeline without the magic: chunk the material, embed the chunks, store them, retrieve the nearest ones, stuff them into the prompt, generate. The retrieval is the part that fails. Chunking too big or too small. An embedding space that cannot tell your two products apart. A vector database that is just a drawer with a better index. Grounding is the claim that the answer came from the pages you handed it, and you can check. NVIDIA&apos;s retriever sits in that drawer. Hallucination is what happens when the book is the wrong book, or no book, and the model talks anyway.

Takeaway: if you cannot point at the page, it was not open book.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Open book. The model does not have to remember your documents. It has to find the right page and read it at the moment you ask. That is retrieval augmented generation, and the honest version is smaller than the demo.

Rob walks the pipeline without the magic: chunk the material, embed the chunks, store them, retrieve the nearest ones, stuff them into the prompt, generate. The retrieval is the part that fails. Chunking too big or too small. An embedding space that cannot tell your two products apart. A vector database that is just a drawer with a better index. Grounding is the claim that the answer came from the pages you handed it, and you can check. NVIDIA&apos;s retriever sits in that drawer. Hallucination is what happens when the book is the wrong book, or no book, and the model talks anyway.

Takeaway: if you cannot point at the page, it was not open book.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Open book. The model does not have to remember your documents. It has to find the right page and read it at the moment you ask. That is retrieval augmented generation, and the honest version is smaller than the demo.</p><p>Rob walks the pipeline without the magic: chunk the material, embed the chunks, store them, retrieve the nearest ones, stuff them into the prompt, generate. The retrieval is the part that fails. Chunking too big or too small. An embedding space that cannot tell your two products apart. A vector database that is just a drawer with a better index. Grounding is the claim that the answer came from the pages you handed it, and you can check. NVIDIA's retriever sits in that drawer. Hallucination is what happens when the book is the wrong book, or no book, and the model talks anyway.</p><p>Takeaway: if you cannot point at the page, it was not open book.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E09. Which Number Means Good: BLEU, ROUGE, Perplexity, FID</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e09</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-09.mp3" length="24846887" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:16:00 +0000</pubDate>
<itunes:duration>997</itunes:duration>
<itunes:episode>9</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>evaluation metrics, BLEU, ROUGE, perplexity, FID, LLM as a judge</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-09.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Which number means good. People quote one figure as if it were a verdict. It is a score on one test, pointed at one kind of mistake.

Rob puts the shelf in order. BLEU for translation, which is overlap with a reference and is said blue. ROUGE for summaries. Perplexity for how surprised the model is by the next word, and lower is better. FID for whether generated pictures look like real ones, lower again. FVD for video. WER for speech, which this series says as word error rate so nobody hears were. CER the letter version. MOS is people listening, higher is better. LPIPS, SSIM, PSNR for pictures. Then the modern awkward one: another model marking the work. Six of those numbers run backwards. Everything else today is higher is better, and the list gets turned around on people constantly.

Takeaway: ask which test, and which way the number is supposed to go.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Which number means good. People quote one figure as if it were a verdict. It is a score on one test, pointed at one kind of mistake.

Rob puts the shelf in order. BLEU for translation, which is overlap with a reference and is said blue. ROUGE for summaries. Perplexity for how surprised the model is by the next word, and lower is better. FID for whether generated pictures look like real ones, lower again. FVD for video. WER for speech, which this series says as word error rate so nobody hears were. CER the letter version. MOS is people listening, higher is better. LPIPS, SSIM, PSNR for pictures. Then the modern awkward one: another model marking the work. Six of those numbers run backwards. Everything else today is higher is better, and the list gets turned around on people constantly.

Takeaway: ask which test, and which way the number is supposed to go.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Which number means good. People quote one figure as if it were a verdict. It is a score on one test, pointed at one kind of mistake.</p><p>Rob puts the shelf in order. BLEU for translation, which is overlap with a reference and is said blue. ROUGE for summaries. Perplexity for how surprised the model is by the next word, and lower is better. FID for whether generated pictures look like real ones, lower again. FVD for video. WER for speech, which this series says as word error rate so nobody hears were. CER the letter version. MOS is people listening, higher is better. LPIPS, SSIM, PSNR for pictures. Then the modern awkward one: another model marking the work. Six of those numbers run backwards. Everything else today is higher is better, and the list gets turned around on people constantly.</p><p>Takeaway: ask which test, and which way the number is supposed to go.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E10. More Than Text: Modalities, CLIP, and One Shared Space</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e10</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-10.mp3" length="23261239" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:18:00 +0000</pubDate>
<itunes:duration>931</itunes:duration>
<itunes:episode>10</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>multimodal AI, CLIP, contrastive learning, zero-shot classification, cross-modal retrieval, SigLIP</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-10.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time.

A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died.

Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm.

Takeaway: meaning often lives between two types, not inside either one.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time.

A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died.

Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm.

Takeaway: meaning often lives between two types, not inside either one.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Somebody says nice job, flat and slow, looking straight past you. The words are positive and the meaning is the opposite, and no machine handed only the words can catch that. Sarcasm lives in the mismatch. That is why more than one type of data has to be in the same model at the same time.</p><p>A modality is a type: words, pictures, sound, video. Multimodal means one model holding more than one of them at once, not two models in a row. A cascade that transcribes then chats is not that. The transcript is where the tone died.</p><p>Then the model that did it from the ground up. CLIP: four hundred million pairs of a picture and the sentence that happened to sit next to it. Two readers, one space, so a picture and a sentence can be compared. Zero-shot classification without a fixed list of labels. SigLIP is the sibling. The grid of in and out is the rest of the hour: captioning, visual question answering, speech in and words out, words in and pictures out, a camera and an instruction and a robot arm.</p><p>Takeaway: meaning often lives between two types, not inside either one.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E11. Patches Are Tokens: The Vision Transformer</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e11</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-11.mp3" length="24454632" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:20:00 +0000</pubDate>
<itunes:duration>981</itunes:duration>
<itunes:episode>11</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>vision transformer, ViT, image patches, inductive bias, CNN vs ViT, self-attention</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-11.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

A picture is a grid of dots. The old way of reading it was a small window sliding around. The new way cuts the picture into squares and feeds the squares in like words. That is the vision transformer. ViT.

Rob walks why that was a genuine surprise: a shape built for sentences, with no built-in sense that nearby pixels belong together, beating the window on pictures once the data is big enough. Patches are tokens. A special token sits at the front and becomes the summary. Position has to be added on purpose, same as in language, because attention is blind to layout. Then the complaints: it wants more data than the window did, it is expensive at high resolution, and a later design called Swin puts hierarchy back in so the window is not the only way to be local.

Takeaway: the same architecture came for pictures, and the square is the word.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

A picture is a grid of dots. The old way of reading it was a small window sliding around. The new way cuts the picture into squares and feeds the squares in like words. That is the vision transformer. ViT.

Rob walks why that was a genuine surprise: a shape built for sentences, with no built-in sense that nearby pixels belong together, beating the window on pictures once the data is big enough. Patches are tokens. A special token sits at the front and becomes the summary. Position has to be added on purpose, same as in language, because attention is blind to layout. Then the complaints: it wants more data than the window did, it is expensive at high resolution, and a later design called Swin puts hierarchy back in so the window is not the only way to be local.

Takeaway: the same architecture came for pictures, and the square is the word.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>A picture is a grid of dots. The old way of reading it was a small window sliding around. The new way cuts the picture into squares and feeds the squares in like words. That is the vision transformer. ViT.</p><p>Rob walks why that was a genuine surprise: a shape built for sentences, with no built-in sense that nearby pixels belong together, beating the window on pictures once the data is big enough. Patches are tokens. A special token sits at the front and becomes the summary. Position has to be added on purpose, same as in language, because attention is blind to layout. Then the complaints: it wants more data than the window did, it is expensive at high resolution, and a later design called Swin puts hierarchy back in so the window is not the only way to be local.</p><p>Takeaway: the same architecture came for pictures, and the square is the word.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E12. Sand and Statues: How Diffusion Makes a Picture</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e12</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-12.mp3" length="24288246" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:22:00 +0000</pubDate>
<itunes:duration>974</itunes:duration>
<itunes:episode>12</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>diffusion models, denoising, classifier-free guidance, latent diffusion, diffusion transformer, text to image</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-12.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Sand and statues. You take a picture and add noise until it is sand. You train a model to take the noise away, one grain at a time, until a statue is standing there. That is diffusion.

Rob keeps it in that picture. The model is a sand detector, not a painter with a plan. Classifier-free guidance is how a sentence steers the sand without a separate classifier. Latent diffusion does the work in the recipe instead of on every pixel, which is why it ran on ordinary hardware. The old backbone was drawn like the letter U. The new one is a diffusion transformer, DiT, and the argument is that it improves more predictably as it grows. Text to image is the demo. Video is the same idea with a clock.

Takeaway: it is not drawing. it is taking noise away on purpose.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Sand and statues. You take a picture and add noise until it is sand. You train a model to take the noise away, one grain at a time, until a statue is standing there. That is diffusion.

Rob keeps it in that picture. The model is a sand detector, not a painter with a plan. Classifier-free guidance is how a sentence steers the sand without a separate classifier. Latent diffusion does the work in the recipe instead of on every pixel, which is why it ran on ordinary hardware. The old backbone was drawn like the letter U. The new one is a diffusion transformer, DiT, and the argument is that it improves more predictably as it grows. Text to image is the demo. Video is the same idea with a clock.

Takeaway: it is not drawing. it is taking noise away on purpose.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Sand and statues. You take a picture and add noise until it is sand. You train a model to take the noise away, one grain at a time, until a statue is standing there. That is diffusion.</p><p>Rob keeps it in that picture. The model is a sand detector, not a painter with a plan. Classifier-free guidance is how a sentence steers the sand without a separate classifier. Latent diffusion does the work in the recipe instead of on every pixel, which is why it ran on ordinary hardware. The old backbone was drawn like the letter U. The new one is a diffusion transformer, DiT, and the argument is that it improves more predictably as it grows. Text to image is the demo. Video is the same idea with a clock.</p><p>Takeaway: it is not drawing. it is taking noise away on purpose.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E13. Pixel-Perfect: U-Net and the Shape of a Medical Problem</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e13</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-13.mp3" length="23736625" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:24:00 +0000</pubDate>
<itunes:duration>951</itunes:duration>
<itunes:episode>13</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>U-Net, image segmentation, skip connections, medical imaging, federated learning, NVIDIA Clara</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-13.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Pixel-perfect. Drawing round a tumour, a cell, an organ, the outline has to land on the right pixels or the drawing is theatre. The shape that is superb at that is named for the letter U.

Skip connections are the whole trick: the fine detail from the way down gets handed across to the way up, so the rebuild is not a blur of the original. Medical images are scarce, private, and labelled by people who are expensive. Federated learning is how several hospitals train without putting the scans in one pile. MONAI is the open medical kit. Parabricks for the genome side. Clara was the umbrella name; the front of the page now lists the pieces. This hour is the medical problem, not a product tour.

Takeaway: the outline is the job, and the skip is why the U works.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Pixel-perfect. Drawing round a tumour, a cell, an organ, the outline has to land on the right pixels or the drawing is theatre. The shape that is superb at that is named for the letter U.

Skip connections are the whole trick: the fine detail from the way down gets handed across to the way up, so the rebuild is not a blur of the original. Medical images are scarce, private, and labelled by people who are expensive. Federated learning is how several hospitals train without putting the scans in one pile. MONAI is the open medical kit. Parabricks for the genome side. Clara was the umbrella name; the front of the page now lists the pieces. This hour is the medical problem, not a product tour.

Takeaway: the outline is the job, and the skip is why the U works.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Pixel-perfect. Drawing round a tumour, a cell, an organ, the outline has to land on the right pixels or the drawing is theatre. The shape that is superb at that is named for the letter U.</p><p>Skip connections are the whole trick: the fine detail from the way down gets handed across to the way up, so the rebuild is not a blur of the original. Medical images are scarce, private, and labelled by people who are expensive. Federated learning is how several hospitals train without putting the scans in one pile. MONAI is the open medical kit. Parabricks for the genome side. Clara was the umbrella name; the front of the page now lists the pieces. This hour is the medical problem, not a product tour.</p><p>Takeaway: the outline is the job, and the skip is why the U works.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E14. Fusion: Eyes, Ears, and Radar</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e14</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-14.mp3" length="24812317" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:26:00 +0000</pubDate>
<itunes:duration>996</itunes:duration>
<itunes:episode>14</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>sensor fusion, early late intermediate fusion, missing modalities, autonomous vehicles, lidar, radar</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-14.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

A car, a robot, a person in a room: the useful picture is never one sensor. Cameras see surfaces. Radar sees velocity through weather. Lidar sees distance. Microphones hear. Fusion is how those disagreeing witnesses become one account.

Rob names three timings. Early: smash the raw signals together before anybody has understood them. Late: each sensor decides, then you vote. Intermediate: each sensor gets part-way, then you combine the part-way. Missing modalities are the real world: fog, a dead camera, a cheap robot that never had lidar. The model that only ever trained with every sensor present will fail the day one is gone. Autonomous vehicles are the worked example. The honest sentence is that fusion is a bet about when the disagreement should be resolved.

Takeaway: one sensor is a witness. several is a case conference.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

A car, a robot, a person in a room: the useful picture is never one sensor. Cameras see surfaces. Radar sees velocity through weather. Lidar sees distance. Microphones hear. Fusion is how those disagreeing witnesses become one account.

Rob names three timings. Early: smash the raw signals together before anybody has understood them. Late: each sensor decides, then you vote. Intermediate: each sensor gets part-way, then you combine the part-way. Missing modalities are the real world: fog, a dead camera, a cheap robot that never had lidar. The model that only ever trained with every sensor present will fail the day one is gone. Autonomous vehicles are the worked example. The honest sentence is that fusion is a bet about when the disagreement should be resolved.

Takeaway: one sensor is a witness. several is a case conference.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>A car, a robot, a person in a room: the useful picture is never one sensor. Cameras see surfaces. Radar sees velocity through weather. Lidar sees distance. Microphones hear. Fusion is how those disagreeing witnesses become one account.</p><p>Rob names three timings. Early: smash the raw signals together before anybody has understood them. Late: each sensor decides, then you vote. Intermediate: each sensor gets part-way, then you combine the part-way. Missing modalities are the real world: fog, a dead camera, a cheap robot that never had lidar. The model that only ever trained with every sensor present will fail the day one is gone. Autonomous vehicles are the worked example. The honest sentence is that fusion is a bet about when the disagreement should be resolved.</p><p>Takeaway: one sensor is a witness. several is a case conference.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E15. Bolting an Eye onto a Language Model: How VLMs Are Built</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e15</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-15.mp3" length="22826747" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:28:00 +0000</pubDate>
<itunes:duration>913</itunes:duration>
<itunes:episode>15</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>vision language models, VLM, projector, visual question answering, visual grounding, NVILA</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-15.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

A language model has never seen a picture. Bolting an eye onto it is a real engineering job, not a metaphor. A vision language model is that bolt: a picture reader, a projector that turns what the reader saw into something the language model can attend to, and then the language model talking.

Rob keeps the assembly in order. Frozen versus trained. Where the projector sits. Visual question answering: an image plus an ordinary question, answered in ordinary language. Grounding: not just describing the picture but pointing at the bit you meant. NVILA is built for machines with no room and no patience. VILA came first. VADER is video, the odd thing that happened in it, and why. Captioning is the easy demo. The hard one is the model that has to be right about a small region.

Takeaway: the eye is a reader plus a translator into the model&apos;s language.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

A language model has never seen a picture. Bolting an eye onto it is a real engineering job, not a metaphor. A vision language model is that bolt: a picture reader, a projector that turns what the reader saw into something the language model can attend to, and then the language model talking.

Rob keeps the assembly in order. Frozen versus trained. Where the projector sits. Visual question answering: an image plus an ordinary question, answered in ordinary language. Grounding: not just describing the picture but pointing at the bit you meant. NVILA is built for machines with no room and no patience. VILA came first. VADER is video, the odd thing that happened in it, and why. Captioning is the easy demo. The hard one is the model that has to be right about a small region.

Takeaway: the eye is a reader plus a translator into the model&apos;s language.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>A language model has never seen a picture. Bolting an eye onto it is a real engineering job, not a metaphor. A vision language model is that bolt: a picture reader, a projector that turns what the reader saw into something the language model can attend to, and then the language model talking.</p><p>Rob keeps the assembly in order. Frozen versus trained. Where the projector sits. Visual question answering: an image plus an ordinary question, answered in ordinary language. Grounding: not just describing the picture but pointing at the bit you meant. NVILA is built for machines with no room and no patience. VILA came first. VADER is video, the odd thing that happened in it, and why. Captioning is the easy demo. The hard one is the model that has to be right about a small region.</p><p>Takeaway: the eye is a reader plus a translator into the model's language.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E16. The Metal: Cores, Bandwidth, Precision, and the Compiler</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e16</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-16.mp3" length="24492866" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:30:00 +0000</pubDate>
<itunes:duration>982</itunes:duration>
<itunes:episode>16</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>CUDA, GPU memory bandwidth, mixed precision, FP8, NVFP4, TensorRT</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-16.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

The metal. Cores, the wires between them, how many bits you spend on each number, and the compiler that turns a model into something those cores will actually run.

CUDA is the programming layer, said like barracuda, not letter-spelled. Memory bandwidth is often the wall, not the arithmetic: the cores are waiting on the next slab of numbers. Mixed precision is using fewer bits where you can afford it. FP8, then a still-coarser format NVIDIA calls NVFP. TensorRT is the compiler: tensor then the letters RT. Hopper, then Blackwell, then Blackwell Ultra, then Rubin paired with a main processor called Vera. Jetson is the same foreman in a module that sits inside a robot or a camera. NGC is the catalogue of already-built pieces.

Takeaway: the bill is often the wire, not the multiply.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

The metal. Cores, the wires between them, how many bits you spend on each number, and the compiler that turns a model into something those cores will actually run.

CUDA is the programming layer, said like barracuda, not letter-spelled. Memory bandwidth is often the wall, not the arithmetic: the cores are waiting on the next slab of numbers. Mixed precision is using fewer bits where you can afford it. FP8, then a still-coarser format NVIDIA calls NVFP. TensorRT is the compiler: tensor then the letters RT. Hopper, then Blackwell, then Blackwell Ultra, then Rubin paired with a main processor called Vera. Jetson is the same foreman in a module that sits inside a robot or a camera. NGC is the catalogue of already-built pieces.

Takeaway: the bill is often the wire, not the multiply.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>The metal. Cores, the wires between them, how many bits you spend on each number, and the compiler that turns a model into something those cores will actually run.</p><p>CUDA is the programming layer, said like barracuda, not letter-spelled. Memory bandwidth is often the wall, not the arithmetic: the cores are waiting on the next slab of numbers. Mixed precision is using fewer bits where you can afford it. FP8, then a still-coarser format NVIDIA calls NVFP. TensorRT is the compiler: tensor then the letters RT. Hopper, then Blackwell, then Blackwell Ultra, then Rubin paired with a main processor called Vera. Jetson is the same foreman in a module that sits inside a robot or a camera. NGC is the catalogue of already-built pieces.</p><p>Takeaway: the bill is often the wire, not the multiply.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E17. Serving It: Dynamo, Dynamo-Triton, and NIM</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e17</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-17.mp3" length="24403167" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:32:00 +0000</pubDate>
<itunes:duration>979</itunes:duration>
<itunes:episode>17</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>model serving, NVIDIA Dynamo, Triton Inference Server, NIM microservices, dynamic batching, KV cache</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-17.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

A finished model sitting on a disk is not a product. Somebody has to stand at the door and answer requests. Serving is that door.

Dynamo is the newer platform for very large language models spread across a great many machines. Dynamo Triton is the older general-purpose server folded into it, and both names are still all over the documentation. NIM is the parcel: a container with the model, the server, and the defaults, so somebody else can run it without building the door. Dynamic batching is waiting a few milliseconds so several questions share the same pass. The KV cache is the notes from earlier in the conversation so you do not redo the whole page every time a new word arrives. When somebody says the model is in production, ask which of those they actually mean.

Takeaway: the parcel is the product. the weights are an ingredient.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

A finished model sitting on a disk is not a product. Somebody has to stand at the door and answer requests. Serving is that door.

Dynamo is the newer platform for very large language models spread across a great many machines. Dynamo Triton is the older general-purpose server folded into it, and both names are still all over the documentation. NIM is the parcel: a container with the model, the server, and the defaults, so somebody else can run it without building the door. Dynamic batching is waiting a few milliseconds so several questions share the same pass. The KV cache is the notes from earlier in the conversation so you do not redo the whole page every time a new word arrives. When somebody says the model is in production, ask which of those they actually mean.

Takeaway: the parcel is the product. the weights are an ingredient.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>A finished model sitting on a disk is not a product. Somebody has to stand at the door and answer requests. Serving is that door.</p><p>Dynamo is the newer platform for very large language models spread across a great many machines. Dynamo Triton is the older general-purpose server folded into it, and both names are still all over the documentation. NIM is the parcel: a container with the model, the server, and the defaults, so somebody else can run it without building the door. Dynamic batching is waiting a few milliseconds so several questions share the same pass. The KV cache is the notes from earlier in the conversation so you do not redo the whole page every time a new word arrives. When somebody says the model is in production, ask which of those they actually mean.</p><p>Takeaway: the parcel is the product. the weights are an ingredient.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E18. The Kitchen: NeMo from Raw Data to a Trained Model</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e18</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-18.mp3" length="24115287" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:34:00 +0000</pubDate>
<itunes:duration>967</itunes:duration>
<itunes:episode>18</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>NVIDIA NeMo, data curation, synthetic data, model parallelism, Riva, TAO Toolkit</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-18.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Last hour talked about a sealed parcel. This hour goes through the door behind it. NeMo is the kitchen: a framework for preparing data, training, adjusting, and measuring, across language, pictures, and speech. It is not a serving layer. It is not a model. You cannot fetch it and ask it a question.

Curator is first on the line: duplicates, quality, personal details, language. Synthetic data is manufacturing material you could not collect. Megatron Core is the library whose job is one training run across a great many machines. Customizer is the stop where you adjust it on your own material. Evaluator runs the benchmarks and the awkward modern one where another model marks the work. Then Dynamo to serve it, NIM to parcel it, Guardrails at the moment it speaks. Riva is speech fast enough to hold a conversation. Metropolis and DeepStream for many video streams at once. TAO is the older toolkit for a smaller custom model, letter-spelled so it is not Taoism.

Takeaway: NeMo is the room, not the dish.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Last hour talked about a sealed parcel. This hour goes through the door behind it. NeMo is the kitchen: a framework for preparing data, training, adjusting, and measuring, across language, pictures, and speech. It is not a serving layer. It is not a model. You cannot fetch it and ask it a question.

Curator is first on the line: duplicates, quality, personal details, language. Synthetic data is manufacturing material you could not collect. Megatron Core is the library whose job is one training run across a great many machines. Customizer is the stop where you adjust it on your own material. Evaluator runs the benchmarks and the awkward modern one where another model marks the work. Then Dynamo to serve it, NIM to parcel it, Guardrails at the moment it speaks. Riva is speech fast enough to hold a conversation. Metropolis and DeepStream for many video streams at once. TAO is the older toolkit for a smaller custom model, letter-spelled so it is not Taoism.

Takeaway: NeMo is the room, not the dish.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Last hour talked about a sealed parcel. This hour goes through the door behind it. NeMo is the kitchen: a framework for preparing data, training, adjusting, and measuring, across language, pictures, and speech. It is not a serving layer. It is not a model. You cannot fetch it and ask it a question.</p><p>Curator is first on the line: duplicates, quality, personal details, language. Synthetic data is manufacturing material you could not collect. Megatron Core is the library whose job is one training run across a great many machines. Customizer is the stop where you adjust it on your own material. Evaluator runs the benchmarks and the awkward modern one where another model marks the work. Then Dynamo to serve it, NIM to parcel it, Guardrails at the moment it speaks. Riva is speech fast enough to hold a conversation. Metropolis and DeepStream for many video streams at once. TAO is the older toolkit for a smaller custom model, letter-spelled so it is not Taoism.</p><p>Takeaway: NeMo is the room, not the dish.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E19. Agents, and the Price of Thinking Longer</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e19</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-19.mp3" length="23749269" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:36:00 +0000</pubDate>
<itunes:duration>951</itunes:duration>
<itunes:episode>19</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>AI agents, function calling, Model Context Protocol, reasoning models, test-time scaling, NeMo Agent Toolkit</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-19.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

An agent is a model that does not only answer. It uses tools. It can look something up, call a function, try again. The price of thinking longer is that every extra step is another chance to be confidently wrong, and another bill.

Rob separates the demo from the loop. Function calling is a structured request, not a vibe. MCP is an open standard for how the thing doing the reasoning finds out what tools exist and calls them, and it is the most confidently misused set of letters in this field. Reasoning models spend test-time compute: they think longer on hard questions instead of only being bigger. That is test-time scaling. The Agent Toolkit is NVIDIA&apos;s kit for wiring this together. Guardrails belong to the next hour. The honest sentence is that an agent is a loop with a bill, not a personality.

Takeaway: thinking longer is a cost decision, not a personality.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

An agent is a model that does not only answer. It uses tools. It can look something up, call a function, try again. The price of thinking longer is that every extra step is another chance to be confidently wrong, and another bill.

Rob separates the demo from the loop. Function calling is a structured request, not a vibe. MCP is an open standard for how the thing doing the reasoning finds out what tools exist and calls them, and it is the most confidently misused set of letters in this field. Reasoning models spend test-time compute: they think longer on hard questions instead of only being bigger. That is test-time scaling. The Agent Toolkit is NVIDIA&apos;s kit for wiring this together. Guardrails belong to the next hour. The honest sentence is that an agent is a loop with a bill, not a personality.

Takeaway: thinking longer is a cost decision, not a personality.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>An agent is a model that does not only answer. It uses tools. It can look something up, call a function, try again. The price of thinking longer is that every extra step is another chance to be confidently wrong, and another bill.</p><p>Rob separates the demo from the loop. Function calling is a structured request, not a vibe. MCP is an open standard for how the thing doing the reasoning finds out what tools exist and calls them, and it is the most confidently misused set of letters in this field. Reasoning models spend test-time compute: they think longer on hard questions instead of only being bigger. That is test-time scaling. The Agent Toolkit is NVIDIA's kit for wiring this together. Guardrails belong to the next hour. The honest sentence is that an agent is a loop with a bill, not a personality.</p><p>Takeaway: thinking longer is a cost decision, not a personality.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E20. Trust: Bias, Guardrails, and Red Teams</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e20</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-20.mp3" length="24863741" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:38:00 +0000</pubDate>
<itunes:duration>998</itunes:duration>
<itunes:episode>20</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>AI bias, trustworthy AI, NeMo Guardrails, prompt injection, explainability, red teaming</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-20.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Trust is not a feeling the model has. It is work you do around it. Bias in the data becomes bias in the answers. A guardrail is a check at the moment it speaks, not a personality transplant. A red team is people trying to make it fail on purpose.

Rob keeps the stops separate. Prompt injection: the user, or the page the model was told to read, smuggles in new instructions. NeMo Guardrails sit at the moment of speech. Explainability is SHAP and LIME, two ways of asking which input pushed the answer, and they are not a window into the model&apos;s mind. A policy is a written rule. A red team is whether the rule survives contact. The episode will not sell you a safe model. It will tell you what the work is called.

Takeaway: rails are at the door. they are not the character of the house.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Trust is not a feeling the model has. It is work you do around it. Bias in the data becomes bias in the answers. A guardrail is a check at the moment it speaks, not a personality transplant. A red team is people trying to make it fail on purpose.

Rob keeps the stops separate. Prompt injection: the user, or the page the model was told to read, smuggles in new instructions. NeMo Guardrails sit at the moment of speech. Explainability is SHAP and LIME, two ways of asking which input pushed the answer, and they are not a window into the model&apos;s mind. A policy is a written rule. A red team is whether the rule survives contact. The episode will not sell you a safe model. It will tell you what the work is called.

Takeaway: rails are at the door. they are not the character of the house.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Trust is not a feeling the model has. It is work you do around it. Bias in the data becomes bias in the answers. A guardrail is a check at the moment it speaks, not a personality transplant. A red team is people trying to make it fail on purpose.</p><p>Rob keeps the stops separate. Prompt injection: the user, or the page the model was told to read, smuggles in new instructions. NeMo Guardrails sit at the moment of speech. Explainability is SHAP and LIME, two ways of asking which input pushed the answer, and they are not a window into the model's mind. A policy is a written rule. A red team is whether the rule survives contact. The episode will not sell you a safe model. It will tell you what the work is called.</p><p>Takeaway: rails are at the door. they are not the character of the house.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E21. Physical AI: Robots, World Models, and Simulated Worlds</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e21</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-21.mp3" length="24946089" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:40:00 +0000</pubDate>
<itunes:duration>1001</itunes:duration>
<itunes:episode>21</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>physical AI, robotics, world models, NVIDIA Cosmos, Omniverse, Isaac GR00T</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-21.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Physical AI. A model that only writes sentences is not enough once the answer has to be an arm moving, a car not hitting a person, a robot in a warehouse. World models generate what happens next in a place. Simulated worlds are where you practice before the real floor.

Omniverse is the place you assemble a copy of somewhere real, accurate enough that physics behaves. Cosmos is a model, not a place: world foundation models pointed at the physical world, generating video of it. Nemotron is the reasoning family. Cosmos Nemotron absorbed VILA, NVILA, and NVLM. Isaac GR00T is humanoids: camera plus instruction, straight to what the arm does. Alpamayo is the same idea pointed at driving. NeMo is not a model. Nemotron, Cosmos, Isaac GR00T and Alpamayo are models. If you carry one sentence out of this hour into an argument, carry that one.

Takeaway: NeMo is what you do to the thing. Nemotron is a thing.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Physical AI. A model that only writes sentences is not enough once the answer has to be an arm moving, a car not hitting a person, a robot in a warehouse. World models generate what happens next in a place. Simulated worlds are where you practice before the real floor.

Omniverse is the place you assemble a copy of somewhere real, accurate enough that physics behaves. Cosmos is a model, not a place: world foundation models pointed at the physical world, generating video of it. Nemotron is the reasoning family. Cosmos Nemotron absorbed VILA, NVILA, and NVLM. Isaac GR00T is humanoids: camera plus instruction, straight to what the arm does. Alpamayo is the same idea pointed at driving. NeMo is not a model. Nemotron, Cosmos, Isaac GR00T and Alpamayo are models. If you carry one sentence out of this hour into an argument, carry that one.

Takeaway: NeMo is what you do to the thing. Nemotron is a thing.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Physical AI. A model that only writes sentences is not enough once the answer has to be an arm moving, a car not hitting a person, a robot in a warehouse. World models generate what happens next in a place. Simulated worlds are where you practice before the real floor.</p><p>Omniverse is the place you assemble a copy of somewhere real, accurate enough that physics behaves. Cosmos is a model, not a place: world foundation models pointed at the physical world, generating video of it. Nemotron is the reasoning family. Cosmos Nemotron absorbed VILA, NVILA, and NVLM. Isaac GR00T is humanoids: camera plus instruction, straight to what the arm does. Alpamayo is the same idea pointed at driving. NeMo is not a model. Nemotron, Cosmos, Isaac GR00T and Alpamayo are models. If you carry one sentence out of this hour into an argument, carry that one.</p><p>Takeaway: NeMo is what you do to the thing. Nemotron is a thing.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E22. The Flywheel: Why NVIDIA Is Hard to Leave</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e22</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-22.mp3" length="25121131" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:42:00 +0000</pubDate>
<itunes:duration>1008</itunes:duration>
<itunes:episode>22</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>NVIDIA moat, CUDA ecosystem, DRIVE AGX Thor, autonomous driving, data flywheel, AI infrastructure</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-22.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

Why one company is hard to walk away from. A meeting you have sat in: the machine is old, the other supplier is cheaper, and every line your team wrote in four years assumes the first machine. The machine is cheaper. The switch is not.

A flywheel is a heavy disc. You push and nothing happens. Eventually it turns under its own weight. The chips are everywhere. Everything generates material. Material trains better models. Better models need more of the machine. And the software those teams wrote assumes CUDA, which turned twenty. Rob tests the flattering version against a receipt: a carmaker&apos;s privacy notice that names the chip company as a joint controller of what the development vehicles record. Then the limit on that receipt. Cars are the public evidence. The software layer is the advantage worth betting on.

Takeaway: the machine is cheaper. the switch is not.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

Why one company is hard to walk away from. A meeting you have sat in: the machine is old, the other supplier is cheaper, and every line your team wrote in four years assumes the first machine. The machine is cheaper. The switch is not.

A flywheel is a heavy disc. You push and nothing happens. Eventually it turns under its own weight. The chips are everywhere. Everything generates material. Material trains better models. Better models need more of the machine. And the software those teams wrote assumes CUDA, which turned twenty. Rob tests the flattering version against a receipt: a carmaker&apos;s privacy notice that names the chip company as a joint controller of what the development vehicles record. Then the limit on that receipt. Cars are the public evidence. The software layer is the advantage worth betting on.

Takeaway: the machine is cheaper. the switch is not.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>Why one company is hard to walk away from. A meeting you have sat in: the machine is old, the other supplier is cheaper, and every line your team wrote in four years assumes the first machine. The machine is cheaper. The switch is not.</p><p>A flywheel is a heavy disc. You push and nothing happens. Eventually it turns under its own weight. The chips are everywhere. Everything generates material. Material trains better models. Better models need more of the machine. And the software those teams wrote assumes CUDA, which turned twenty. Rob tests the flattering version against a receipt: a carmaker's privacy notice that names the chip company as a joint controller of what the development vehicles record. Then the limit on that receipt. Cars are the public evidence. The software layer is the advantage worth betting on.</p><p>Takeaway: the machine is cheaper. the switch is not.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
<item>
<title>S1E23. What Moved: The Words This Season Earned</title>
<guid isPermaLink="false">cloudadorn.academy.ngai.s01e23</guid>
<enclosure url="https://media.cloudadorn.org/nvidia-generative-ai/audio/episode-23.mp3" length="23757430" type="audio/mpeg"/>
<pubDate>Tue, 25 Aug 2026 18:44:00 +0000</pubDate>
<itunes:duration>952</itunes:duration>
<itunes:episode>23</itunes:episode>
<itunes:season>1</itunes:season>
<itunes:episodeType>full</itunes:episodeType>
<itunes:explicit>false</itunes:explicit>
<itunes:image href="https://media.cloudadorn.org/nvidia-generative-ai/art/show-cover-3000.jpg"/>
<itunes:keywords>generative AI glossary, Dynamo-Triton rename, NeMo vs Nemotron, NVFP4, omnimodal models, season recap</itunes:keywords>
<podcast:transcript url="https://media.cloudadorn.org/nvidia-generative-ai/transcripts/episode-23.txt" type="text/plain"/>
<description>This episode is narrated using AI voice technology. The content and script are original.

The last hour is the words this season earned, said again without the scaffolding. Nesting dolls, attention as a heap, the whisper chain, two opposite training failures, shapes not brands, the cheapest sentence first, the small piece called LoRA, open book only if you can point at the page, which number and which way it runs, one shared space, patches as tokens, sand and statues, the U, the case conference, the bolted-on eye, the wire not the multiply, the parcel not the weights, the kitchen that is not a model, the agent as a loop with a bill, rails at the door, physical AI, the flywheel.

Names that moved while we were talking: Dynamo Triton, NeMo versus Nemotron, NVFP, omnimodal. CLIP still generates nothing and is still a foundation model. When a product says AI, it has told you almost nothing. Ask which doll.

Takeaway: carry the words, not the product names.

Subscribe for the rest of the season, and visit cloudadorn.com.</description>
<itunes:summary>This episode is narrated using AI voice technology. The content and script are original.

The last hour is the words this season earned, said again without the scaffolding. Nesting dolls, attention as a heap, the whisper chain, two opposite training failures, shapes not brands, the cheapest sentence first, the small piece called LoRA, open book only if you can point at the page, which number and which way it runs, one shared space, patches as tokens, sand and statues, the U, the case conference, the bolted-on eye, the wire not the multiply, the parcel not the weights, the kitchen that is not a model, the agent as a loop with a bill, rails at the door, physical AI, the flywheel.

Names that moved while we were talking: Dynamo Triton, NeMo versus Nemotron, NVFP, omnimodal. CLIP still generates nothing and is still a foundation model. When a product says AI, it has told you almost nothing. Ask which doll.

Takeaway: carry the words, not the product names.

Subscribe for the rest of the season, and visit cloudadorn.com.</itunes:summary>
<content:encoded><![CDATA[<p>This episode is narrated using AI voice technology. The content and script are original.</p><p>The last hour is the words this season earned, said again without the scaffolding. Nesting dolls, attention as a heap, the whisper chain, two opposite training failures, shapes not brands, the cheapest sentence first, the small piece called LoRA, open book only if you can point at the page, which number and which way it runs, one shared space, patches as tokens, sand and statues, the U, the case conference, the bolted-on eye, the wire not the multiply, the parcel not the weights, the kitchen that is not a model, the agent as a loop with a bill, rails at the door, physical AI, the flywheel.</p><p>Names that moved while we were talking: Dynamo Triton, NeMo versus Nemotron, NVFP, omnimodal. CLIP still generates nothing and is still a foundation model. When a product says AI, it has told you almost nothing. Ask which doll.</p><p>Takeaway: carry the words, not the product names.</p><p>Subscribe for the rest of the season, and visit cloudadorn.com.</p>]]></content:encoded>
</item>
</channel></rss>
