Your Premiums Live in a Table. Your Exclusions Live in a PDF. Neither Survives Retrieval Well.
Why do AI assistants quote your cover correctly but miss the exclusion, and what happens to a premium written as five lakh?

An insurance page is built out of three things a retrieval system handles poorly. The premium sits in a table cell. The comparison sits in an image. The exclusion that makes the coverage statement accurate sits twelve pages into a policy document, separated from the clause it modifies.
Then the whole thing gets chunked into passages, and a passage is what the assistant actually reads.
The result is a specific and under-discussed failure: the answer a customer receives is not usually wrong about whether you cover something. It is wrong about the conditions under which you cover it, because the condition and the coverage got separated somewhere between your page and the response.
What finance pages are actually made of
We pulled the served HTML of fourteen insurance, lending and investing product pages across India, the US and the UK on 23 September 2026, and counted how they encode their facts.
Indian pages lean on tables. American pages do not.
| Page | Tables in served HTML | Lakh or crore references | Rupee symbols |
|---|---|---|---|
| Policybazaar, term insurance | 13 | 128 | 85 |
| HDFC ERGO, health | 5 | 45 | 50 |
| BankBazaar, home loan | 4 | 50 | 0 |
| Acko, car | 3 | 19 | 21 |
| GEICO, auto | 0 | 0 | 0 |
| Progressive, auto | 0 | 0 | 0 |
| Policygenius, life | 0 | 0 | 0 |
The American pages in our sample served no tables at all. The Indian pages served between three and thirteen, and they carry the numbers a buyer wants.
This is not an argument against tables. A comparison grid is the right way to present nine plan variants to a human. It does mean Indian BFSI has a larger share of its commercially important facts sitting in a structure that behaves differently from prose when it is chunked and embedded.
The same amount, written three ways
Look at the last two columns. BankBazaar's home loan page carries fifty references to lakh or crore and not a single rupee character.
So a sum appears as "50 lakh", elsewhere as `₹50,00,000`, elsewhere as "5000000", and in a table cell possibly as "50L". To an Indian reader these are one number. To a parser they are four distinct strings, and the Indian grouping convention puts the separators in positions the rest of the world does not use.
Any firm relying on a retrieval system to quote its figures accurately in the Indian market has a normalisation problem that American and British competitors simply do not have. We have not seen anybody write about this, and it affects every insurer and lender in the country.
The mechanism is easier to see than to describe. The diagram below shows a policy document splitting into passages, and which one gets retrieved.

Why chunking is the part that bites
The assistant never reads your page
Retrieval systems split documents into passages, embed the passages, and rank them against the question. What comes back to the model is a handful of passages, not your document.
Every assumption your page makes about what the reader has already seen is an assumption the passage cannot carry.
Where a policy document breaks
Policy wordings are the most cross-referential documents in commercial writing. A coverage clause in section 3 is qualified by a definition in section 1 and an exclusion in section 8. That structure is deliberate and legally sound.
Chunk it and the coverage clause becomes a self-contained passage asserting that something is covered. The exclusion becomes a separate passage, competing for retrieval on its own merits, and it frequently loses because the question the customer asked matched the coverage language.
The output is an answer that quotes your document accurately and describes your product wrongly. That is a harder failure to catch than an outright fabrication, because every individual sentence checks out.
A worked example
Take a health policy that covers day-care procedures, subject to a twenty-four month waiting period for pre-existing conditions.
Your page says so, correctly, in two places. The features section states that day-care procedures are covered. The waiting periods section, further down under its own heading, states the twenty-four month condition.
A customer asks an assistant whether the policy covers a day-care procedure for a condition they already have. The retrieval system matches the question against the features passage, which says day-care procedures are covered and says nothing about waiting periods. The waiting-period passage does not match as strongly, because the customer did not use the phrase.
The assistant answers yes.
Nothing was fabricated. The passage was quoted faithfully. The customer now expects cover they will not receive for two years, and the first anyone hears about it is at claim time.
Where a table breaks
A table's meaning lives in the relationship between a cell and its headers. Lift the cell out and it is a number with no subject.
Some systems preserve that structure. Some flatten it. The variable most under your control is whether the number exists anywhere else on the page in a sentence that names what it is.
Where an image breaks
A rate card published as a PNG carries its numbers in pixels. Whether a given system extracts them is a per-vendor question, and alt text long enough to carry a full rate table is not realistic.
If a figure exists only inside an image, treat it as unavailable for quotation.
Why the stakes are higher in a regulated market
In an unregulated category, a retrieval failure costs you a click.
In insurance and lending it produces a customer with a documented expectation that your policy does not meet, formed from an answer that quoted your own published material. Whether that has regulatory implications where you operate is a question for your compliance team, and it is worth them knowing that it can happen at all. In our experience most compliance functions have never been told that their carefully qualified disclosures can be split from the claims they qualify.
The disclosure regimes in India and the UK both assume a reader encountering a communication whole. A chunked passage is not that, and nobody drafting those rules was thinking about embedding models.
What to do about it
Restate every table in a sentence
Not the whole table. The two or three figures that carry the decision.
If the grid shows premiums for nine plan and age combinations, add a line of prose naming the typical case: cover of a stated amount for a stated age at a stated premium. The table stays for humans. The sentence is what gets quoted.
Write conditions into the same passage as the claim
This is the highest-value change on the list and the one most likely to meet resistance.
Wherever your marketing pages state coverage, state the principal condition in the same paragraph. Not a link to the exclusions page. Not a footnote marker. The same passage, so that a chunk containing the claim also contains the qualifier.
Compliance teams usually like this more than marketing does, which makes it an easier internal sell than it first appears.
Normalise your numbers, at least once
For Indian sites, state key amounts in more than one form where it reads naturally. "Cover of 50 lakh (₹50,00,000)" costs nothing and removes the ambiguity for a parser that has only ever seen Western grouping.
Pick one canonical form for structured data and use it consistently.
Give the important numbers a home in text
Any figure a customer might ask an assistant about should appear once, on your site, in server-rendered prose, in a sentence that names it. Premiums, limits, deductibles, fees, eligibility thresholds.
A number that only exists in a calculator, a table cell or an image is a number you have declined to publish in a quotable form, which is the same rendering question approached from a different angle.
Publish the exclusions as web pages
The exclusions are usually the most-asked and least-published part of any policy. They live in a PDF because that is where the legal text lives.
A plainly written exclusions page, one per product, in HTML, with each exclusion stated as a self-contained sentence, is one of the few pieces of content in this category that is both genuinely useful to customers and structurally suited to retrieval.
Use product structured data for the figures that matter
Where a page describes a specific product with specific terms, express the key attributes in structured data rather than relying on the prose alone.
Structured data is unambiguous in a way prose is not: a stated value with a stated currency does not need a parser to guess at lakh notation or a rupee glyph. It will not carry your whole rate card, and it does not need to. It needs to carry the handful of figures you would least like to be misquoted on.
Validate whatever you ship. Malformed markup is not a partial win, it is ignored.
Test what comes back
Ask each assistant a specific question about your own product. Does this policy cover X. What is the premium for a 35-year-old. What is excluded.
Compare the answer against your documents. Where it is wrong, find the passage the wrong answer came from and fix the chunk rather than the page. Where it invents, that is the brand misinformation problem, which has its own remedy.
Re-read your own pages as isolated passages
A cheap exercise that finds most of these problems. Take your top product page, split it at every heading, and read each section as though it were the only thing you had.
Ask of each block: does this make a claim whose qualifier lives somewhere else? Would somebody acting on this block alone be misled?
Anywhere the answer is yes, you have found a chunk that can be retrieved on its own and produce a wrong impression. Most product pages yield three or four in under half an hour.
The short version
The facts that decide an insurance purchase are stored in the three formats least suited to being retrieved as passages, and the exclusion is the one that gets orphaned most reliably.
Restate the key numbers in prose, put conditions next to claims, and publish your exclusions as pages rather than as clauses buried in a document. None of that requires new product, and it fixes the failure that matters most: being described as more generous than you actually are, by a system quoting you accurately.
We work through this in insurance and financial services, and the structural side sits with technical GEO and SEO.



