The same patient appears in training and validation. Does that matter?

Answer it out loud before you open anything. The value of the flags below is in comparing them to what you actually said — including whether you asked about the data before naming a model.

The production scenario behind the question

A healthcare startup classifies skin lesion images. Each patient contributed several photos of the same lesions over time. The random split by image gives a strong validation result; a pilot at a new clinic, with new patients, performs much worse. The team suspects the clinic's cameras.

What it is really testing

Whether the candidate can recognise entity leakage from the setup alone, and knows that the split must match the way the model will be used — on *new* patients. A good answer also handles the harder version: what if the product is used on returning patients, and does the split change?

Where the mechanism is taught