An overall score can hide a gap, so accuracy can be broken down by group. Disaggregating TS-40 shows SH-1 got 26 of 30 right in the three classes taught in English but only 4 of 10 in Spanish. NIST names three kinds of AI bias, and each can happen without anyone meaning to be unfair. A generative model can work worse for some languages, and people may wrongly trust it to work equally well for everyone. Generative AI can also be too uniform (homogenization), and training too much on AI-made data can lead to model collapse. Fixing bias does not by itself make a system fair. The crew rates risk R2, harmful bias and homogenization.