NVIDIA Generative AI. When Training Goes Wrong: Overfitting, Bad Data, and Metrics That Lie. Hosts: Shelley and Rob. Cloudadorn Academy. This episode is narrated using AI voice technology. The content and script are original. Shelley: I'm Shelley, I ask the questions. Rob builds these things for a living, which is why he gets the answers. Rob: Morning. Shelley: I want to talk about failure. Somebody trains one of these things, it comes out bad, and they come to you. What do you ask? Rob: Bad how. Shelley: Bad is bad. Rob: It really isn't. There are two ways for this to go wrong, they're opposites, and the fix for one makes the other worse. Shelley: But you can tell them apart. Rob: In about ten seconds, with the right two numbers in front of you. Shelley: Then give me the two numbers. Rob: Before the numbers, two drivers. Shelley: Drivers, as in cars. Rob: The first learned on exactly one route. The route to work. They know every pothole, they know which corner is tighter than it looks. Shelley: That sounds good. Rob: On that route they're flawless. Put them on any other street and they're helpless. They didn't learn to drive. They learned a road. Shelley: And the second one? Rob: Had one lesson and stopped. Lost on their own route, lost everywhere else, equally bad wherever you put them. And the one you actually want is a third driver, who learned to drive. Shelley: So how do you find out which one you've got? Rob: By putting them on a road they have never seen. There's no other way. Shelley: And with a model that means what? Rob: Before you train anything, you take a slice of your data and lock it in a drawer. People call that slice the validation set, or the held out set. I'm going to call it the drawer. The model never learns from it. It's the street it's never driven. Shelley: So I've got two numbers. How it does on what it studied, and on the drawer. Rob: Watching that pair move is the whole diagnosis. First pattern. The number on what it studied keeps getting better. The number on the drawer gets worse. Shelley: Better on one, worse on the other. Rob: That's the first driver. It's memorising the route, and the name for it's overfitting. It learned your data rather than the pattern in it, including the parts that are just noise. Shelley: And the second pattern? Rob: Both numbers are bad and both stay bad. It never got good at the thing it studied either. Shelley: That's the one lesson. Rob: That's underfitting. Either the model is too simple to hold the pattern, or you didn't train it long enough. Shelley: And from a distance those two look the same. Rob: From a distance they're both a disappointing number in an email. Which is exactly the problem. Shelley: So do the fixes. You said they were opposite. Rob: Take the memoriser first. Six things help, in two piles. Shelley: Start with pile one. Rob: Give it more world. More data, or more examples made out of the ones you have, changed in ways that stay honest. Rotate the photograph. Crop it. It's still a cat. Shelley: And that counts? Rob: It counts as a new example. That's augmentation, and the point is to stop it seeing the identical thing until it has that memorised rather than the idea. Shelley: And pile two? Rob: Give it less room to memorise in. Make the model smaller. Or leave it the same size and hold it back while it learns. Shelley: Hold it back how? Rob: Switch bits of it off at random as it goes, so nothing can rely on anything else being there. And charge it a standing fee for making any of its numbers large. Shelley: And the sixth? Rob: The sixth is stopping earlier, and it needs its own few minutes. Shelley: Then the one who had a single lesson. Rob: Everything reverses. Bigger model, not smaller. More to look at in each example, not less. Train longer, not shorter. And ease off everything you were holding it back with. Shelley: The same actions pointing the other way. Rob: Exactly that. And here's the trap, which catches good people. Shelley: Go on, what's the trap. Rob: Both problems arrive as a bad number. And a bad number triggers one instinct in everybody. We need more data. Shelley: Which is sometimes right? Rob: For the memorising problem it's the strongest thing you can do. Better than a cleverer method. Shelley: And for the other one? Rob: Nothing at all. If it's doing badly on the examples already in front of it, it isn't using what it has. Shelley: So somebody spends six months collecting data. Rob: Six months, real money, and the number hasn't moved, because they treated the wrong disease. Shelley: Unpark stopping earlier. Rob: You're training. Every pass over your data you check both numbers and write them down. Shelley: A pass being one full trip through everything you've got. Rob: One full trip, which people call an epoch. Now draw it. Time across the bottom, badness up the side, two lines. Shelley: And they do what? Rob: The line for the data it studied, its loss, goes down and keeps going down forever. Shelley: Of course it does. It's memorising. Rob: The other line, the one from the drawer, comes down with it for a while. Then it flattens. Then it turns and climbs. Shelley: And the turn is where it stopped learning and started memorising. Rob: That's the picture. The most useful chart in this work. Shelley: So you stop there. Rob: You stop there and you keep the version from just before the turn. Not the one you ended up with. Shelley: And that's early stopping? Rob: That's early stopping. Now the honest part, because that was the version in the diagram. Shelley: As opposed to? Rob: Reality. That second line isn't a smooth curve. It's bumpy. Up, then down again, then up. Shelley: So if you stop the first time it ticks upward. Rob: You may have stopped on a wobble and thrown away a better model that was three passes away. Shelley: Then what do you actually do? Rob: Wait. Give it a fixed number of passes with no improvement before you believe it. And keep the best version you've seen. Shelley: So the rule isn't stop when it turns. Rob: The rule is stop when it has stopped improving, and be patient about deciding that. Less tidy, and the true one. Shelley: All of this is after you've trained something. Rob: Which is late. The best work happens before any of it. Shelley: Doing what, exactly? Rob: Looking at your data. Actually looking. How the values are spread. What's missing? What's weirdly extreme? Whether your categories are anywhere near even. And whether the answer has got into the questions. Shelley: Is that a real activity, or you being tidy? Rob: It has a name and a long history. Exploratory data analysis. Mostly pictures, done before you model anything, to get surprised early rather than late. Shelley: Make the case. Rob: It's the part everybody skips. So. Some researchers took ten of the collections this industry measures itself against and checked the labels. Shelley: Checked whether the answers were right. Rob: They found errors in every single one. In the big image collection, more than one label in twenty was simply wrong. Shelley: One in twenty. On the thing everyone is scored against. Rob: And in a study of people building these systems for serious purposes, around nine in ten had a project damaged by a data problem that was upstream, invisible and avoidable. Shelley: Nine in ten. Rob: So when somebody asks whether to fix the data or tune the model, the honest answer, an enormous amount of the time, is fix the data. Shelley: Then tell me what to look for. Rob: Five things, each with a symptom you can recognise. First, lopsided categories. You're hunting something rare, and ninety nine out of a hundred cases are nothing. Shelley: And the model learns to say nothing. Rob: It learns that guessing nothing is right nearly every time, which it's. Its overall number looks superb. It has never found the thing you built it for. Shelley: What do you do? Rob: Rebalance what you feed it, or tell it the rare cases count for more. And stop judging it on the overall number, which we'll come to. Shelley: Second thing wrong. Rob: Gaps. Values that simply aren't there. Shelley: So fill them in. Rob: You can, from what you know. But there's a nicer move people forget, which is to add a column recording that it was missing, because the gap itself is often information. Shelley: Third thing wrong. Rob: Extremes. One value miles from everything else. It wrecks your averages and makes training lurch about. Shelley: So throw it out. Rob: That's the instinct and it's wrong. Find out why it's there first. A broken sensor, a typing mistake. Shelley: And sometimes it isn't a mistake? Rob: Sometimes the strange one is what you were looking for. If you're hunting fraud, the outlier is the entire job. Shelley: Fourth thing wrong. Rob: Leakage. This is the nasty one. Something got into what the model learned from that it won't have on the day it has to answer for real. Shelley: Give me one. Rob: You're predicting which customers will cancel, and one of your columns is the date they were sent a cancellation confirmation. Shelley: Ah. It isn't predicting. It's reading the answer. Rob: And it will be magnificent. Practically perfect. Which is the symptom. Results too good to be true, right until it meets real work and falls over. Shelley: So throw out that column. Rob: Throw out anything that secretly contains the answer. Then a second fix, which catches far more people. Split your data before you do anything else to it. Shelley: Anything else like what? Rob: Say you want everything on a common scale, so you take the average across all your data and divide by it. Shelley: That seems harmless. Rob: That average now contains the drawer. What you locked away has already influenced what you fed the model, and your measurement isn't honest. Shelley: So split first. Rob: Split first, always. It's one line of work and one of the commonest ways to fool yourself. Shelley: And the fifth? Rob: Duplicates. The same example appearing twice. Shelley: Surely that's just redundant. Rob: Not if it appears in what the model studied and also in the drawer. Then your drawer isn't a road it has never driven, and your measurement is inflated by an amount you can't know. Shelley: How bad is this really? Rob: In one standard collection scraped off the open web, there's a single English sentence, about sixty words long, that appears more than sixty thousand times. Shelley: Sixty thousand times. Rob: So it isn't a rounding error. Shelley: You keep saying the overall number. What are the numbers actually? Rob: For sorting things into categories, four are worth knowing. Start with the obvious one. What fraction of its answers were correct. That's accuracy, and it's perfectly fine when your categories are roughly even. Shelley: And when they're not? Rob: Then it's worse than useless, because it's confidently wrong. Ninety nine out of a hundred cases are nothing, remember. Shelley: And a model that says nothing every time. Rob: Is ninety nine percent accurate. Put that on a slide and it looks like a triumph. It has found precisely zero of the things it was built to find. Shelley: So accuracy is out. Rob: For lopsided work accuracy is essentially never the answer. Which is why the next two exist. Shelley: Give me the next two. Rob: Of the ones you flagged, how many were actually right. That's precision. Shelley: And the other half? Rob: Of all the real ones out there, how many did you catch. That's recall. Shelley: They sound like the same thing said twice. Rob: They point at two different mistakes. Precision is about false alarms. Recall is about misses. Which one you want depends on which costs more. Shelley: Give me one. Rob: Filtering unwanted mail. What's the expensive mistake? Shelley: Throwing away something real. I'd rather see junk than lose a letter. Rob: So you care about precision. When it flags something it had better be right, because a false alarm is a letter in the bin. Shelley: And the opposite? Rob: Screening a population for a disease. Now the expensive mistake is missing somebody who has it. Shelley: So you want recall. Catch everyone, accept some false alarms. Rob: Because a false alarm costs a second look, and a miss costs a person. Shelley: Can you cheat either of them? Rob: Both, trivially, which is why you never quote one alone. Want perfect precision? Flag almost nothing. Find the one case you're certain about, and you were right every time you spoke. Shelley: And perfect recall? Rob: Say yes to everybody. You have caught every real case in the population. You have also flagged the entire population. Shelley: So each alone is gameable. Rob: Which brings the fourth. Combine them so a weak half drags the whole thing down. Shelley: So you can't hide a bad one behind a good one. Rob: You can't. It's called F one, and it's the usual default for lopsided work. Both halves have to be decent before it looks decent. Shelley: Then that's the answer. Rob: It has one weakness. One number, so it won't tell you which half is the weak one. For that you look at them separately, or at the picture of the trade between them. Shelley: And where do all four come from? Rob: One small table. Sort every answer four ways. Said yes and was right. Said yes and was wrong. Said no and was wrong. Said no and was right. Shelley: Four boxes, then. Rob: Called the confusion matrix, which is a horrible name for a simple object. Shelley: And the four numbers live in there. Rob: All four. Though which line is truth and which is the guess depends on whose convention you're reading, and the people who write this software say so themselves. Shelley: Do the chart part properly. Rob: Happily, because it's genuinely easy. You match the picture to the shape of the thing. Shelley: Meaning what, in practice. Rob: Never which chart is nicest. It's how many things am I looking at, and are they categories or a range. Shelley: Start with the one you showed me. Rob: A line, for something changing along a continuous run. Ours was badness against passes, two lines together, and that's the picture of memorising. Shelley: What's the next chart? Rob: Bars, for comparing across things that are separate. How many examples in each category? Or one bar per version of the model. Shelley: So bars for categories, lines for a range. Rob: That distinction is most of the skill. Next. One measurement across all your examples, and you want its shape. Where it sits, how spread out, does it lean, are there strays at the edge. A histogram. Shelley: And with two measurements? Rob: A scatter. One dot per example, one measurement across, one up. It answers whether the two are related at all. What the model predicted against what actually happened, say. Shelley: What else is there? Rob: A box plot, for when you care about the spread and the tail rather than the average. The middle of the data drawn as a box, the extremes hanging off it. Shelley: When would I want that? Rob: How long your system takes to answer, across a lot of requests. You want to see the slow ones. Shelley: And a grid of colours? Rob: A heatmap. Colour stands in for size, across two dimensions. And you have already met the thing you draw with it. Shelley: The four boxes. Rob: Drawn as colour, so which mistake dominates is visible from across the room. Shelley: Give me the last one. Rob: The circle cut into slices. Parts of one whole. What your data is made of, in proportions. Shelley: That's the one everybody makes fun of. Rob: People have opinions. What it's for is proportions of a single whole, and that much is true. Shelley: Today in one breath. Rob: Two failures, opposites. It memorised your data, or it never learned it. You can't tell which without a slice you locked away first. The fixes point opposite ways, so treating the wrong one costs months. Shelley: And before any of that? Rob: Look at your data. They checked ten of the collections this whole field is scored against and found wrong labels in all ten. And when you judge the result, the overall fraction correct is the number that lies to you most often. Shelley: The bit that got me was that it took somebody going and looking. Rob: Which is why the unglamorous work wins. Nobody gets excited about checking a spreadsheet. It moves more than the clever idea does. Shelley: One more. You keep saying smaller model, bigger model. Like there's a shelf of them. Rob: There's a shelf. A long one, and everything on it has one thing it's superb at and one it's hopeless at. Shelley: Then that's next, and I want the whole shelf.