An accessory list contains “USB-C plug”, “USB-A plug” and simply “cable”. Ask an assistant to put every entry into USB-C or USB-A and the last row invites a guess. A complete-looking table can conceal incomplete information.
Define what a label is allowed to mean
For this example, the task is to identify the single connector type explicitly named in each description. Add Unknown when no type is stated, and Needs review when the description names multiple types or contradicts itself. The two specific plugs have clear labels; “cable” remains Unknown. “USB-A to USB-C cable” needs review under this single-label rule, even though its description is informative.
Ask for three columns: the original entry, the assigned category and the words that support the choice. Preserve a record identifier so you can trace each result back to its source. A plausible explanation is not enough if those supporting words are absent.
Review confident answers as well
Keep a small set of examples with expected answers and compare results whenever you change the categories or instructions. Inspect a sample of confident labels too. An assistant's certainty is not a measured accuracy score.
NIST describes how generative systems can confidently produce false content or diverge from supplied information. The connector exercise is our practical illustration of that risk, not a documented product test. An Unknown option makes missing evidence visible; it does not ensure the model will use it correctly.
Keep unresolved rows in the result and collect the missing details. Removing them only to make the table look finished hides useful work that still needs doing.