Exam Room · Advanced Generative AI Developer

Content Moderation With Rekognition, Comprehend, and Guardrails

· 35 min read

Generative AI Development · part of The Exam Room

The situation

A community app has bolted a generative feature onto an existing user-generated-content platform, and the surface area for unsafe content has grown with it. Members upload profile photos and short clips. They record voice notes that the app plays back to other members. They write posts and comments in free text. On top of that, a new assistant powered by a model on Amazon Bedrock takes a member’s prompt and generates replies, captions, and summaries that get shown to everyone else.

Every one of those paths can carry something the platform should not publish: explicit or violent imagery, a slur buried in a comment, a phone number or credit-card detail pasted into a post, a voice note that is abusive, or a model completion that drifts into a topic the brand has said it will never discuss. The team’s first instinct was to reach for one moderation API and run everything through it. That does not exist. Images are not text, audio is not an image, and moderating what a member typed is a different problem from moderating what the model returned.

The real question is a routing question. For each kind of content, and at each stage of the pipeline, which AWS service is built to screen it, and how do those services combine into one pipeline that covers uploads, prompts, and outputs without gaps.

What actually matters

The first thing that decides everything is modality. A moderation service is trained on one kind of media and takes no input in the others. Amazon Rekognition analyses pixels and returns nothing about the meaning of a sentence. Amazon Comprehend analyses text and accepts no image. Audio is a third case: no text or vision service takes an audio file, so it has to be turned into text first. Getting the modality match right is most of the decision, and forcing one service onto the wrong media is the classic mistake.

The second axis is the stage in the pipeline, which matters most once a generative model is in the loop. There are three distinct places content needs screening, and they are not interchangeable. User uploads arrive before the model and are screened at ingestion. The prompt going into the model is an input that can carry abuse or an attempt to steer the model somewhere unsafe. The completion coming out of the model is fresh content the model just produced, and it needs screening before it reaches another member even if the prompt was clean. A tool aimed at uploads does nothing for what the model generates, and vice versa.

Third is what “unsafe” even means here, because it is several different concerns, not one. Explicit imagery is one thing; personally identifiable information like a phone number or a card is a completely different detection problem; toxicity and harassment is a third; and staying off brand-forbidden topics is a fourth. Some services cover one of these, some cover several, and the ones that overlap do so at different stages, so knowing which concern you are solving narrows the field fast.

Fourth is confidence and the grey zone. None of these services returns a clean yes or no; they return labels with confidence scores, and the platform sets the threshold. That immediately creates a band of borderline cases that sit below the auto-block line but above the auto-approve line, and those are exactly the ones a human should look at. A moderation design that has no path for the uncertain middle either over-blocks safe content or ships unsafe content, so a human-review step for the borderline band is part of the architecture, not an afterthought.

Fifth, these are building blocks rather than a finished moderation product, and the coverage comes from the pipeline. A single upload might need Rekognition on the image and Comprehend on the caption; a voice note needs Transcribe then Comprehend; a model turn needs Guardrails on both ends. The services are designed to be composed, and the moderation posture comes from wiring the right ones into each path rather than from any one call.

What we’ll filter on

  1. Modality, is the content an image, video, audio, or text?
  2. Pipeline stage, is this a user upload, a model input, or a model output?
  3. Concern, explicit and unsafe visuals, PII, toxicity, or a brand-forbidden topic?
  4. Confidence handling, does the path route the borderline band to human review?
  5. Composability, does the content need two services chained rather than one?

The landscape

Amazon Rekognition content moderation. The vision service for still images and stored video; there is no content-moderation path for a live stream. Version 7 of the moderation model returns labels on a three-level taxonomy. Top-level categories include Explicit, Violence, Visually Disturbing, Drugs & Tobacco, Alcohol, Gambling, Rude Gestures, and Hate Symbols. Each label carries a confidence score and the level it sits at. AWS recommends filtering on level one or two and reserving level three for concepts you want to exempt. For images you call DetectModerationLabels synchronously, or StartMediaAnalysisJob for a bulk asynchronous run; for stored video you start StartContentModeration and collect the results with GetContentModeration, which timestamps where in the clip each label appears. It reads pixels only, so a caption attached to the image is out of its scope.

Amazon Comprehend for text. The natural-language service, and it covers more than one moderation concern. Its PII detection works on English or Spanish and finds entities like names, phone numbers, email addresses, and payment-card numbers in free text. Locating them is available in real time; redacting them in place is an asynchronous batch job only, so a synchronous path takes the entity types and character offsets back and does the masking itself. Toxicity detection is a separate synchronous call, DetectToxicContent, which takes English text as a list of segments. It returns a confidence score on each of seven labels, profanity, hate speech, insult, graphic, harassment or abuse, sexual, and violence or threat, plus an overall toxicity score for the segment. The request limits shape the integration. A call carries at most ten segments, 1 KB each and 10 KB across the list, so a long post, or the transcript of a several-minute voice note, gets chunked before it can be screened. Chunk on sentence boundaries; scores taken on half-sentences wobble. Per-label scores let the application set its own numeric threshold per category, strict on hate speech and looser on profanity, where a guardrail content filter offers four strength levels per category and no raw score.

Comprehend also trains bespoke safety classifiers. Some unwanted content is specific to this platform and matches no general category: a coded slur that means nothing outside this community, a recruitment scam that keeps resurfacing in the same shape. Custom classification learns those from the team’s own labelled prompts and posts, single-label or multi-label. Deployed behind a real-time endpoint it sits in the request path, validating content before it is stored, instead of a batch job reading content after it has been published. Comprehend is the reader for posts, comments, and any text extracted from another modality. It takes no image input, so abusive text inside a screenshot never reaches it; that path is a Rekognition DetectText call, which returns up to 100 words per image, feeding its output onward.

Amazon Transcribe as the audio bridge. No vision or text service takes an audio file, so audio is converted to text first, and Transcribe is that step. It also carries its own moderation options. A vocabulary filter masks, removes, or tags a supplied list of words, in batch and streaming jobs alike. Toxicity detection scores speech segments across seven categories, using acoustic cues as well as the words. It runs on batch transcriptions in US English only, so a live stream or a member speaking another language gets no signal from it. In practice you transcribe the audio, act on Transcribe’s toxicity output where it is available, and pass the resulting transcript into Comprehend for the fuller PII and toxicity read. Audio is never moderated directly; it is transcribed, then the text is moderated.

Amazon Bedrock Guardrails. The moderation layer for the generative model itself, and the one that sits on the prompt and the completion rather than on stored uploads. A guardrail bundles several policies. Content filters cover hate, insults, sexual content, violence, misconduct, and prompt attacks, each at None, Low, Medium, or High strength. Denied topics run to thirty per guardrail, each a name plus a definition of at most 200 characters; a match returns the blocked message you configured instead of the completion. Word and profanity filters handle exact terms. Sensitive-information policies block or mask PII on the prompt and on the response. Content filters read images as well as text, for PNG and JPEG files up to 4 MB and twenty images a request. A guardrail therefore screens a photo a member attaches to a prompt, and an image a model generates. The sensitive-information filter is text-only. A guardrail is applied through the ApplyGuardrail API directly or by attaching it to a Bedrock model invocation. It covers what a member asked and what the model returned. An upload that sits in a bucket and never reaches the model is outside it.

Human review for the grey zone. The layer that catches everything that is neither a confident block nor a confident approve. A workflow holds the borderline item, presents it to a reviewer, and feeds the verdict back, with a sample of confident calls pulled in for quality auditing. Amazon Augmented AI (A2I) shipped this workflow ready-made, with a worker task template and a direct integration with Rekognition content moderation, and it keeps running for teams already on it; it closed to new customers in late July 2026. A fresh build assembles the same loop from primitives: a Step Functions workflow or an SQS queue holding the flagged item, a reviewer UI you own, and a callback that resumes the pipeline with the decision.

Evaluation

Side by side

Service Modality Stage it screens PII Toxicity Visual unsafe content Denied topics
Rekognition content moderation Image, stored video User uploads ✗ ✗ ✓ ✗
Comprehend Text Uploads and extracted text ✓ ✓ ✗ ✗ (custom classifier approximates)
Transcribe Audio to text User uploads (audio) ✗ (masks via vocab filter) ✓ (batch, en-US) ✗ ✗
Bedrock Guardrails Text and image (model I/O) Model input and output ✓ (text only) ✓ ✓ (at the model boundary) ✓
Human review Any (review) Borderline band n/a n/a n/a n/a

Reading the table against the app: uploaded photos and clips go to Rekognition, posts and comments go to Comprehend, and voice notes go through Transcribe then Comprehend. The assistant’s prompts and completions go through Guardrails on both ends. Anything in the uncertain middle of those calls goes to human review.

The moderation map

Content Screened by Image or video upload profile photo, clip Rekognition moderation labels, confidence Text upload post, comment Comprehend PII, toxicity, custom classifier Audio upload voice note Transcribe audio to text Model prompt what the member asked Bedrock Guardrails filters, denied topics, PII Model completion what the model generated Human review borderline confidence band

The solution

Rekognition for the uploaded pixels. Point it at the image or the stored video and it returns moderation labels on the three-level taxonomy, each with a confidence score. The platform sets a MinConfidence on the request, which defaults to 50%, and a block threshold on the results, so you decide how aggressive to be; a dating app and a children’s app draw the line in different places. For video the job is asynchronous and the results are timestamped, which lets you flag the exact second an issue appears rather than rejecting the whole clip blind. Rekognition also has text-in-image detection, which matters when abuse arrives as a screenshot; you pull the text out and hand it to Comprehend, because the moderation labels themselves are about visual content, not the words printed on it.

Comprehend for every path that ends in text. Its two moderation-relevant features solve different concerns. PII detection is an entity problem, locating names, numbers, and card details, and it is what stops a member publishing someone’s phone number. In a real-time path it hands back entity types and offsets and the application does the masking, because Comprehend’s own redaction is an asynchronous batch job. Toxicity detection is a classification problem, scoring text for harassment, hate, threats, and profanity. When the unwanted content is specific to this community and not a generic category, a custom classifier trained on the platform’s own labelled examples fills the gap. Comprehend is also the second half of the audio and screenshot paths, reading text that a different service extracted.

Comprehend has a second job in the generative path as well. Running it over a member’s text before that text reaches the model gives the layered stack its outer ring. Comprehend screens and masks on the way in. A guardrail sits on the prompt and the completion at the model boundary, and a Lambda post-processing check validates what comes back before the app renders it. That is the same defence in depth the injection layers are built from, so the moderation routing here and that stack describe one architecture. Comprehend adds a ring outside the boundary; it does not take the boundary over, because the guardrail is the only layer here that evaluates the completion.

Transcribe as the only way audio gets moderated. There is no direct audio-moderation service in this stack, so the pattern is fixed: transcribe first, then read the transcript. Transcribe adds two moderation options of its own. A vocabulary filter masks, removes, or tags a supplied word list at transcription time. Toxicity detection scores speech segments using acoustic cues as well as the words, which catches tone a plain transcript loses. That second one is batch-only and US English only, so it is not available on every path. Either way the fuller PII and toxicity read comes from passing the transcript to Comprehend, so audio is a two-service chain by design.

Guardrails for both ends of the model. This is the pick people miss, because uploads and generation feel like the same “moderation” job but they are not. A guardrail is attached to the model invocation or called through ApplyGuardrail, and it screens the prompt going in and the completion coming out. Content filters catch hate, insults, sexual content, violence, misconduct, and prompt attacks at strengths you set, across text and attached or generated images. Denied topics let the brand describe, in plain language, subjects the assistant should stay off; a match returns the blocked message rather than the completion, and no upload-facing service offers that. The sensitive-information policy blocks or masks PII on either side, on text only. Choosing which model sits behind the guardrail is a related exercise. Bedrock’s automatic model evaluation carries toxicity as a built-in metric alongside accuracy and robustness, scored against the RealToxicityPrompts and BOLD datasets, so candidates can be compared on how often they emit unsafe output before the guardrail sees it. Screening the output matters even when the input was clean, because the completion is new content the model just produced, and it is the thing that actually gets shown to another member.

Human review for the band nobody can auto-decide. Every service here returns confidence, not certainty, so set two thresholds rather than one: above the upper line, auto-block; below the lower line, auto-approve; in between, route to a human-review workflow. For a team already running A2I, that workflow exists, with a worker task template and a direct Rekognition integration. A new build queues the flagged item, surfaces it in its own reviewer UI, and writes the verdict back into the pipeline. Sampling a slice of the confident decisions through the same review loop is how you catch threshold drift before members do.

Worked example

A member submits a single post that carries three things at once: a photo, a typed caption, and a voice note, and then asks the assistant to write a summary of it for the feed.

The photo goes to Rekognition. DetectModerationLabels comes back with “Violence” at 91% and the platform’s block threshold is 80%, so the image is held. Because it is over the block line, it does not need a human; if it had come back at, say, 74%, it would fall in the review band and go to a human reviewer instead of being published on a guess.

The caption goes to Comprehend. Toxicity detection scores it low, but PII detection returns a phone number with its offsets, so the application masks that span before the caption is stored rather than rejecting the whole post. One modality, two different concerns, one service handling both.

No text service takes the voice note as it stands, so Transcribe converts it. Its vocabulary filter masks a couple of slurs inline, and, because this is a batch job in US English, its toxicity scores flag one segment as harassment. The transcript then goes to Comprehend for the same PII and toxicity read as the caption got, because the audio path always ends in a text service.

Finally the member asks the assistant to summarise the post, and that model turn is wrapped in a Bedrock guardrail. The prompt is screened on the way in. The generated summary is screened on the way out against the content filters and denied topics, so even a clean prompt cannot produce a completion that reopens the violent content or drifts onto a forbidden subject. Four services, one post, each piece routed to the tool built for its media and its stage, and the human-review loop taking whatever lands in the middle.

What’s worth remembering

  1. Modality decides the service first: Rekognition for images and stored video, Comprehend for text, and audio has to be transcribed before any text tool can read it.
  2. Screen model output as well as input, because the completion is fresh content the model produced and can be unsafe even when the prompt was clean.
  3. Uploads, model inputs, and model outputs are three separate stages; a tool built for one does nothing for the others.
  4. Every service returns confidence, not a verdict, so set an auto-block and an auto-approve threshold and route the band between them to human review.
  5. Real moderation is a pipeline, not a call: a single post can need Rekognition, Comprehend, Transcribe, and Guardrails together, each on the piece it was built for.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.