Labels arrive 30 days after the prediction. What does that do to your monitoring, and what does the dashboard mean today?
Answer it out loud before you open anything. The value of the flags below is in comparing them to what you actually said — including whether you asked about the data before naming a model.
A lending model's dashboard shows a rolling precision of 100% for the last four weeks and the team has stopped looking at it. Defaults are only known 30 days after the due date, and the dashboard counts every prediction with no label yet as "correct so far". A new model was deployed three weeks ago.
React to this
Say what you would question, what you would trust, and what you would need to know first.
Model health — rolling 28 days (illustrative) predictions: 41,206 labelled: 0 precision: 100.0% (unlabelled counted as correct) recall: n/a approval rate: 58.1% (prev model, same period last year: 52.4%) model version: v9 since 21 days ago; v8 no longer scoring
What it is really testing
Whether the candidate can see that the dashboard's recent window contains no outcomes at all, and can design monitoring that is honest about the delay: outcome metrics reported only for matured cohorts, proxies for the unmatured ones, and the consequence that the new model cannot be judged on outcomes yet.