Q. You must measure a GenAI feature for bias and toxicity before launch. Which service?
A. SageMaker Clarify’s foundation-model evaluation (the fmeval library) scores accuracy, toxicity, semantic robustness, and prompt stereotyping (bias). Bedrock model evaluation also offers toxicity and stereotyping metrics.
Why? Bias for a generative model shows up as stereotyping and quality disparity, measured offline, not as label parity.