Exam Room · AI Business Strategist

The Damage Claim With Forty Photos

· 27 min read

AI for the Business · part of The Exam Room

The situation

A regional home insurer writes about 190,000 household policies across three states. Storm claims arrive in bursts: 290 property claims in an ordinary week, 4,100 in the week after a front comes through.

One file from the June batch holds four things. Forty photographs from the policyholder’s phone, most of them the same two metres of ridge capping from slightly different angles. A twenty-two minute recorded call describing what she heard at two in the morning and what she found at seven. A four-page assessor’s report scanned from paper, handwriting in the margins, a signature at the bottom. And one row in the claims system carrying fifteen fields: policy number, peril, date of loss, cause of loss, reserve, and ten more.

Three uses are queued behind a single funded line of AUD$380,000 for the year. Triage claims by likely severity on the day they arrive, so the worst reach a senior handler. Draft the assessor’s summary for the assessor to edit rather than write cold. Flag which claims are worth sending somebody out to see. The claims director has to say which of the three this year’s material can support. Everyone expects the same answer: the row is usable, the rest is unstructured, and unstructured is the expensive kind.

What actually matters

Artificial intelligence is the umbrella: any system doing work that would otherwise take human judgement. Machine learning gets its behaviour from historical examples rather than written rules, so it needs past cases with the answer attached. Generative AI sits inside machine learning, built on foundation models that produce content from a prompt, with the expensive training already done on somebody else’s material. That distinction sets the timeline on any ask. A capability that learns from this insurer’s history waits on whatever the insurer failed to write down; one that arrives trained reads the photographs today.

Structured data is carved into fields with agreed meanings. The claims row is structured; a photograph is a grid of pixels, a recording a waveform, a scan a picture of a page. The price ranking bolted onto that split was correct when a photograph could only be used by paying somebody to describe it. Foundation models now take an image or a document as input directly, a managed service turns a call into a labelled transcript, and the row everybody trusts carries a cause-of-loss field the storm and escape-of-water teams have filled in with different meanings since 2020.

The split describes shape, not cost. Four questions decide cost. Can a model read the thing as it stands? Must it be pulled into fields first? Does its meaning have to be agreed across systems before a comparison means anything? Did it exist when the decision was made? The photographs answer yes, no, no and yes, which is why they are cheap. Cause of loss is already queryable and fails the third, expensive in a way no storage bill reveals.

Material can clear all four and still be wrong. An incomplete field shrinks the population the answer was fitted to, and nothing in the output says so. An inaccurate value is wrong in the same direction every time. Inconsistent meaning goes unnoticed longest, because both teams fill the field in correctly by their own lights. A reserve refreshed monthly cannot carry a decision taken on the morning a claim arrives. And where the outcome was never recorded, no amount of cleaning creates it.

What we’ll filter on

  1. Readable as it stands. Can a model take this material as input, or does an extraction project have to finish before the use can start?
  2. One agreed meaning where claims get compared. Where the use ranks, routes or scores across claims, does the field it leans on mean the same thing in every row?
  3. Present at decision time. Did the material exist at the moment the decision has to be made, or does it only appear afterwards?
  4. Outcome recorded. Is the thing being predicted written down somewhere in the history, in a form that says what actually happened?
  5. A cost shape that can be priced before the work starts. Per image, per page, per minute, per token, per seat and per instance-hour are all quotable in an afternoon. A labelling bill is not.

The landscape

Forty photographs of a roof

Vision-capable foundation models on Amazon Bedrock take an image as input and answer questions about it, billed on tokens consumed. Amazon Bedrock Data Automation covers the same material as a managed job and bills per image. Neither needs a labelled photograph from this insurer; the training that recognises a roof happened on somebody else’s pictures. Most of the forty show the same ridge capping, so feeding all forty is billed forty times for one roof, and selecting a handful belongs in the design.

A twenty-two minute recording

Bedrock Data Automation processes audio and returns a full transcript and a summary of the call. Speaker labelling and a topic-by-topic summary are not on by default, so they belong in the specification rather than being assumed. English is supported, and AWS publishes a four-hour ceiling on a single audio file, so a twenty-two minute call is nowhere near it. Billing is per minute, a figure finance can multiply by the claim count.

A scanned report with handwriting on it

The same service handles documents, the material written off hardest of the four. Bedrock Data Automation recognises handwritten characters as well as printed ones, and English is one of the six input languages it reads. It attaches confidence scores and visual grounding to every extracted field, so a handler can see where a number came from. Billing is per page, and a four-page report is a rounding error.

Amazon Bedrock Knowledge Bases reads text for nothing extra, and that default returns nothing at all from a photograph or a scanned page. Reading figures, charts, tables and images means paying per page or per token to parse them. The choice is made per source of data, so a photo library and a scanned-report library can differ, and it then applies to every PDF in that source, including the ones that held nothing but text.

Fifteen fields in the claims system

The material the business already trusts is the only one in the file with a defect that stops a use. Amazon Quick is where reporting over it lands, with Amazon Quick Sight for dashboards and Amazon Quick Index for grounding answers in company documents, priced per seat. A reporting layer over a field carrying two meanings publishes the disagreement rather than resolving it, and nothing in the row says which definition a record was written under.

The history nobody wrote down

Amazon SageMaker AI is where a model fitted to this insurer’s own settled claims is built and served, with a real-time endpoint priced per instance-hour whether claims are arriving or not. Settlement value is recorded on every closed claim, so severity has six years of outcomes to learn from. Whether a site visit was worth making is recorded nowhere: the system holds that a visit happened, not whether it changed anything, and nothing about the claims where nobody went. Creating that history means people reading closed files and typing verdicts, a labelling bill nobody has quoted.

Evaluation

Side by side

Candidate use Readable as it stands Agreed meaning where compared Present at decision time Outcome recorded Priceable cost shape
Triage by severity on arrival ✓ ✓ ✓ ✓ ✓
Draft the assessor’s summary ✓ ✓ ✓ ✓ ✓
Flag the claims worth a visit ✓ ✗ ✓ ✗ ✗
Rank on the claims row alone ✓ ✗ ✓ ✓ ✓

Triage takes its tick in the second column on a condition the claims director should hear while the money is allocated: it clears the column only because the design leaves cause of loss out and builds the severity signal from the photographs, the call and the settled value.

Drafting ticks the fourth column because it has no outcome to predict. It composes rather than estimates, so past assessor summaries are house style rather than labels.

The visit flag fails three columns independently: cause of loss routes claims today, so a flag built on it inherits two definitions; the verdict it would reproduce was never recorded; and nobody has sized the labelling exercise that would fix either. The fourth row ticks four of five columns and still loses, because the photographs and the call carry the severity of a storm claim.

Where each ask lands

THE THREE ASKS, ONE BUDGET LINE Triage by likely severity on the day the claim arrives Draft the assessor's summary for the assessor to edit Flag the claims worth sending somebody to see Can a model read the material as it stands? Nothing falls out here now. Photographs, scanned pages and a transcript are all read directly. all three Where claims get compared, does the field mean one thing? Cause of loss: two teams have filled it in with two meanings since 2020. Routing leans on it. yes no Was the outcome you want ever written down? yes Funded this quarter Triage and drafting. Per image, per minute, per page, per token. no Can the gap be named with an owner, a cost and a date? yes Remediate, then re-price Site-visit flag. One new field recorded from Monday, re-priced in a year. no Not a project yet No owner, no cost, no date to fund.
The first gate no longer catches anything, which is the change most business cases have not absorbed. The second and third gates catch the field everybody trusts and the verdict nobody recorded.

The solution

Fund triage and drafting as one stream over one ingestion of the file, not two projects that each parse the same claim.

Bedrock Data Automation runs once per claim over the selected photographs, the recording and the scanned report, producing a transcript with the speakers labelled, extracted fields with confidence scores and visual grounding, and captions for the images. A foundation model on Bedrock reads that output twice, once to place the claim in a severity band and once to draft the summary. The claims row goes in as context and stays out of the severity signal, because cause of loss is under repair.

Three gotchas belong in the design. The per-image charge means somebody decides how many of the forty photographs go in, and “all of them” is billed forty times for one roof. The default knowledge-base parser returns nothing from a photograph or a scanned page, so a retrieval layer configured without checking indexes nothing from those files and raises no error. And a confidence score is an estimate: the threshold above which a value goes straight into the file and below which it reaches a handler is a claims decision, in the same way that the confidence threshold on an automated decision belongs to the business that carries the errors.

The cause-of-loss field gets a named repair: one owner, the claims operations lead; one definition agreed between the storm team and the escape-of-water team and written down; a back-fill over the last eighteen months rather than six years, because the reporting that matters looks at the recent past; a date. Until it lands, nothing routes or ranks on the field, and the reporting seats in Amazon Quick note which definition applies to which period.

The visit flag is deferred with a collection job attached, not refused. The verdict takes one dropdown on the assessor’s closing screen and one line in a procedure: did this visit change the outcome, yes, no, or partly. That is a form change and a fortnight of somebody’s attention. Twelve months later the business has a labelled set covering a full storm season, and twelve is the honest number because a set built out of one quiet quarter teaches a model about quiet quarters.

Re-price the document estate while the money is in the room. This insurer wrote off its scanned material three years ago on advice that was correct then; it is now readable at a per-page rate anybody can multiply out. Re-pricing it takes an afternoon, a sample file and the published rate.

Drafting runs on a vendor’s model that changes each time a new version ships, so a saved set of example files and expected summaries runs against every version before it goes live. Triage drifts as claims patterns change. Both belong in the business case.

Worked example

The forty photographs are readable as they stand, need no extraction, carry no meaning to agree with anyone, and arrived on day one, so both funded uses can have them at a per-image cost set by how many go in. The call is readable once transcribed, a managed job at a per-minute price, and nothing in it gets compared across claims.

The scanned report is readable, handwriting included, at a per-page price, but it exists only after somebody has been sent to look. Drafting happens after the visit and can use it; triage happens before and cannot. The same document, the same quality, two opposite answers, separated by when it existed rather than what it is made of.

The fifteen-field row needs no extraction and is the only material in the file that blocks a use. Cause of loss means one thing to the storm team and another to the escape-of-water team, so anything ranking or routing on it compares incompatible things, and the reserve is refreshed monthly, too stale for a decision taken on the morning the claim lands. What the data already is decides which method is available, and knowing that before the budget meeting matters more than any service comparison.

What’s worth remembering

  1. Shape is not cost. Structured against unstructured describes how material is shaped; four other questions decide what it costs to use.
  2. Four questions set the price. Readable as it stands? Must it be pulled into fields? Meaning agreed across systems? Present at decision time?
  3. Generative AI arrives trained. Photographs, recordings and scanned pages work on day one; a model fitted to this business’s history waits on what nobody recorded.
  4. Check the trusted field. The unstructured material is usually fine; two teams have recorded two meanings in one field for years, neither of them wrong.
  5. Cost shape follows the material. Per image, page, minute, token, seat and instance-hour all quote up front; a labelling bill does not.
  6. Deferred is not refused. A use the material cannot support gets an owner, a cost, one new thing recorded from Monday, and a re-pricing date.

These posts are LLM-aided. Backbone, original writing, and structure by Craig. Research and editing by Craig + LLM. Proof-reading by Craig.