A score only means something if the test matches how the system will really be used. Validation is objective evidence that a system meets the requirements for its intended use, and reliability is working as required, without failure, over time. Accuracy is how close results are to the true values; it should be measured on a clear, realistic test set and include false positive and false negative rates. Good scores on tests made for humans do not prove an AI system fits a job. The crew writes TS-40 from class notes, scores SH-1 at 30 of 40, tests the Check feature with CK-20, and rates the likelihood of risk R1.