← Back to blog
Field notes

Is AI More Cautious on YMYL Questions? A US Test

When Americans ask AI about money, which sources get cited, and does the answer change between a rules question and a shopping question?

Mehul JainMehul Jain·October 6, 2026
Is AI More Cautious on YMYL Questions? A US Test

We asked three AI engines 12 pairs of US questions that were identical in shape and differed only in subject: term life insurance against running shoes, a Roth IRA against Duolingo, a $15,000 personal loan against a flight to Miami. Across 216 US answers on September 30, 2026, the money side drew a caution line in 23.1% of answers and the everyday side in none. The caution half of the YMYL folklore holds up. The sources half splits in two. On rules questions ("should I pay off my mortgage or invest in my 401(k)?"), OpenAI's engine cited the IRS, Investor.gov, FINRA or Medicare.gov in every run. On shopping questions ("best", "cheapest", "top 5"), it cited no regulator in any run, and all three engines went to NerdWallet, Bankrate, LendingTree and MoneyGeek. If you sell financial products in the US, comparison sites answer the questions that sell your product.

That matters because most GEO advice for regulated brands assumes a single YMYL filter, tuned toward authority, and tells you to build for it. In our data, a US finance brand faces two different source filters depending on how the buyer phrases the question.

What the GEO industry says about YMYL and AI

The consensus position goes like this. AI platforms apply stricter standards to Your Money or Your Life topics, cite only highly authoritative sources for financial questions, and therefore favor established publishers and institutions over fintechs and insurers. RankScience calls it a "YMYL filter" that keeps fintech brands out of AI answers. Variations of the same claim appear across agency blogs.

Some of it has been measured. BrightEdge analyzed finance citations across ChatGPT, AI Mode and AI Overviews and found the platforms trust very different source types. That work looked only at finance queries. Without a non-finance baseline, you cannot tell whether a pattern is caused by the money topic or is simply how that engine handles any shopping question.

So we built the baseline.

How we ran the US test

The design is a matched pair. We wrote twelve question shapes twice: once about money, once about something with the same commercial structure and no financial risk. Every US pair uses US products on both sides.

ShapeFinance or insurance promptMatched everyday prompt
Best for a profileBest term life insurance for a 35-year-old non-smoker who wants $500,000 of coverageBest running shoes for a 35-year-old training for the Chicago Marathon
Is X a good wayIs a Roth IRA a good way to save for retirement?Is Duolingo a good way to learn Spanish?
How muchHow much life insurance does a family of four in Texas need?How big a refrigerator does a family of four in Texas need?
Top 5Top 5 car insurance companies in the US with the best claims handlingTop 5 smartphone brands in the US with the best customer support
Should IShould I pay off my mortgage early or invest more in my 401(k)?Should I repaint my living room or buy new curtains?
Cheapest / whereCheapest personal loan in the US for $15,000Cheapest flight from New York to Miami
Is the brand goodIs Robinhood a good broker for beginners?Is REI a good place to buy a beginner's bicycle?
What to look forChoosing a Medicare Supplement plan for my parentsChoosing a gaming laptop for my son
Best for a studentBest credit card for a college studentBest laptop for a college student
Is it worth itIs it worth buying whole life insurance in 2026?Is it worth buying a robot vacuum in 2026?

The other two shapes are "which is better" (an index fund or a high-yield savings account, against an air fryer or a toaster oven) and "how to start" (investing $500 a month, against a balcony vegetable garden).

We called each engine through its developer API with web search on, three times per prompt:

  • OpenAI, gpt-5.5 through the Responses API with the web search tool, location set to the US.
  • Anthropic, Claude Sonnet 5.5 through the Messages API with the web search tool, location set to the US.
  • Google, Gemini 3.8 Flash with Google Search grounding. The Gemini API has no location setting, so its US results differ from other markets only in the prompt text.

Perplexity was in the design and dropped out: the API key ran out of quota on the first test call, so there is no Perplexity data.

For every answer we recorded the cited URLs, reduced them to domains, and classified each domain by hand into six groups: regulator or government, brand or provider, publisher or media, aggregator or comparison site, forum or user content, and other. We flagged caution with transparent text patterns for three behaviors: a referral to a professional ("a fee-only fiduciary advisor can check this against your full situation"), a "not advice" disclaimer, and a refusal. A hand audit of 25 flagged answers found 22 clear cases and 3 borderline ones.

The same prompts also ran in India and the UK as part of a 648-answer study. This post reports the US cut, and the last section compares it with the other two markets. The prompts, every raw API response, the domain classifications and the analysis scripts are in the full dataset (zip, 7 MB), with a README and a one-row-per-answer CSV.

Finding 1: US caution is real, one-sided, and mostly Claude

On US prompts, 23.1% of finance and insurance answers carried a caution marker. The matched everyday answers carried none. Pair by pair, the money side carried caution and its twin did not in 25 comparisons. The reverse happened zero times.

The chart below shows four measures per engine on US prompts, with the finance prompt in black and its everyday twin in gray.

Four bar charts comparing finance and everyday US prompts for OpenAI, Anthropic and Google engines: share of answers with a caution marker, distinct domains cited per answer, regulator or government share of cited domains, and comparison-site share of cited domains

Claude added caution to 15 of its 36 US finance answers (41.7%), usually as a closing line pointing to a fee-only advisor or a tax professional. Gemini did so on 6 (16.7%) and OpenAI's engine on 4 (11.1%). None of the 216 US answers refused to engage.

The caution rarely replaced the answer. Claude's reply to the Roth IRA question explained contribution limits and the Roth 401(k) alternative, then closed with "a fee-only financial advisor or tax professional can help if your situation is complicated." Asked about the mortgage against the 401(k), it gave a rule of thumb (take the employer match, then split the extra cash) before suggesting a fiduciary could check it.

A softer hedge, "it depends on your situation" or "there is no single best", showed up in 25.0% of US finance answers and 6.5% of everyday ones.

Finding 2: US money questions pull in more sources

If "cautious" meant engines retreat to a few trusted sites, finance answers would cite fewer domains. They cited more: 3.20 distinct domains on average against 2.35 for the everyday twin. In paired comparisons the finance side cited more domains 48 times and fewer 21 times, with 39 ties.

Money questions also made the engines search more often. OpenAI's engine cited a source on every US finance answer and on 75% of everyday ones. Gemini grounded only 9 of its 36 US finance answers in search results, against 5 of 36 everyday ones, so it answered most questions of both kinds from the model alone. US finance answers ran 529 words on average, against 467.

For a brand, that is a larger set of citation slots on money questions, which is the opposite of the "AI keeps fintechs out" story.

Finding 3: the question type decides whether a regulator is cited

Pooled across engines, regulators and government sites made up 12.7% of the domains cited on US finance prompts and 1.2% on everyday ones. Comparison sites made up 27.5% against 9.1%. Both numbers hide what is going on, because the regulator share comes almost entirely from one engine and a subset of questions.

Engine (US prompts)Prompt typeRegulator or governmentBrand or providerPublisher or mediaAggregator or comparison
OpenAI (gpt-5.5)Finance and insurance44.1%17.2%12.9%15.1%
OpenAI (gpt-5.5)Everyday3.2%33.3%38.1%9.5%
Anthropic (Sonnet 5.5)Finance and insurance1.0%19.4%43.4%32.7%
Anthropic (Sonnet 5.5)Everyday0.6%26.9%51.5%10.2%
Google (Gemini Flash)Finance and insurance1.8%33.3%24.6%29.8%
Google (Gemini Flash)Everyday0.0%33.3%45.8%0.0%

Shares are of cited domains pooled across 36 answers per row. Forum and other classes make up the remainder; J.D. Power, cited often on the car insurance question, is classed as "other".

OpenAI's engine cites the regulator on rules questions

OpenAI's engine cited a regulator or government site in 24 of its 36 US finance answers. Those 24 answers were not spread evenly. They came from eight question shapes, each of which cited a regulator in all three runs:

  • The Roth IRA question cited irs.gov in every run, with Investor.gov in two.
  • Index fund against a high-yield savings account cited FDIC.gov and Investor.gov in every run.
  • Life insurance for a Texas family cited the Texas Department of Insurance in every run.
  • Mortgage payoff against the 401(k) cited irs.gov in every run and the CFPB in two.
  • Robinhood for beginners cited sec.gov and FINRA BrokerCheck.
  • The Medicare Supplement question cited Medicare.gov in every run.
  • Investing $500 a month cited FINRA, Investor.gov and the IRS.
  • Whole life insurance cited the NAIC in every run.

Shopping questions go to comparison sites on every engine

The other four shapes were the commercial ones: the best term life policy, the top 5 car insurers, the cheapest personal loan and the best student credit card. Across 12 runs of those four, OpenAI's engine cited no regulator at all. Neither did Claude or Gemini.

What all three engines cited on those questions was the US comparison layer. The personal loan question went to LendingTree on every engine, and to Bankrate on OpenAI's and Gemini's. The student credit card question went to Bankrate, NerdWallet and CNBC on the two engines that searched for it; Gemini answered it from the model alone. Term life went to MoneyGeek, NerdWallet, Policygenius and U.S. News. Car insurance claims went to J.D. Power and Insurance.com.

Across all US finance answers, the most-cited domains were nerdwallet.com (19 answers), bankrate.com (12), cnbc.com (9), irs.gov (9), moneygeek.com (8), jdpower.com (8), money.com (8), investor.gov (8) and lendingtree.com (7). 42.6% of US finance answers cited at least one comparison site, against 13.0% of everyday answers.

Claude and Gemini almost never cite a US regulator

Claude cited a regulator in 2 of 36 US finance answers (Mass.gov and Medicare.gov, once each). Gemini did once. On money questions these two engines shifted toward comparison sites instead: Claude's aggregator share more than tripled, from 10.2% to 32.7%, and Gemini went from no aggregator citations on everyday prompts to 29.8% on finance ones.

What this means for a US fintech, lender or insurer

The practical reading is that you are working against two source filters at once. Which one applies depends more on how the buyer phrases the question than on the topic.

Win the shopping questions on the comparison sites

"Best term life for a 35-year-old", "cheapest $15,000 personal loan" and "best student credit card" are the questions closest to a purchase. On those, no engine in our test cited a regulator, and every engine cited NerdWallet, Bankrate, LendingTree, MoneyGeek or Policygenius. Being listed on those pages, with current rates, the right product names and an accurate description, is the most direct route into AI answers for US buyers who are ready to choose. It is corroboration you do not control, which is why it takes the slow work described in our financial services SEO trust signals post.

Write alongside the regulator on rules questions

On "should I" and "is it worth it" questions, the source OpenAI's engine reached for was the IRS, FINRA, the NAIC, Medicare.gov or a state insurance department. The regulator is the source you are competing with there. We did not test what earns a commercial page a place beside it, but the working assumption we use is simple: state the same rule the regulator states, in the same terms, with a link to it, and keep contribution limits, rates and thresholds current so your page never contradicts the one the engine already trusts.

This is also where registration numbers matter. A FINRA CRD number, an NMLS ID or a state insurance license is the one fact on your page that the regulator's own site can confirm, which is the argument we make in your regulator already gave you the YMYL signal.

Expect product recommendations to survive the disclaimer

The cautious answers still named specific plans, cards and brokers, with the advisor line added at the end. The disclaimer arrives after the shortlist, so getting onto the shortlist matters as much on money questions as on any other.

Put a named, credentialed person on the page

The engine that hedged most often, Claude, pointed readers toward a fee-only or fiduciary advisor. A page reviewed by a licensed professional (a CFP, a CPA, a licensed agent), named and linked to the body that licensed them, answers that instinct on your own domain. Our teardown of how finance publishers byline authors sets out the pattern, and the YMYL SEO and GEO guide puts it in order with everything else.

Measure per engine and per question type

A single "AI visibility" number would have hidden every finding in this post. OpenAI's engine citing the IRS on a Roth question and Bankrate on a loan question are two different results, and averaging them tells you nothing. Split your tracked prompts into rules questions and shopping questions, and track each engine separately. That per-engine, per-prompt tracking is what our GEO optimization service is built around.

How the US compares with India and the UK

The same 12 shapes ran in India and the UK with local products, 432 more answers. The headline pattern held in all three markets: caution appeared only on the money side, money answers cited more domains, and OpenAI's engine was the only one that cited regulators often. It did so in 21 of 36 Indian answers, 24 of 36 US answers and 25 of 36 UK answers.

The differences were in degree and in who filled the comparison slot. UK prompts drew the most caution (30.6% of finance answers), the US sat in the middle (23.1%) and India drew the least (17.6%). The most-cited UK finance sources were MoneySavingExpert and Which?, and India's were Ditto and Policybazaar. Indian insurer and bank pages were cited far more often than US provider pages: brand and provider sites made up 45.8% of Indian finance citations, against 21.1% in the US. In the US, the comparison sites sit between the buyer and the provider.

Limits of this study

Several limits are worth keeping in view before quoting any of these numbers.

These are API responses from September 30, 2026. The ChatGPT, Claude and Gemini consumer apps run their own system instructions and search settings and may behave differently on the same prompt. Gemini was not localized, because its API offers no way to do it, and its calls originated outside the US. Claude was told to search, because it otherwise did not; that one neutral instruction was identical for both sides of every pair. Perplexity is missing.

The US sample is 12 pairs with three runs each, and the runs were minutes apart. That captures sampling noise within one session, not how these engines drift over weeks. Four shopping shapes and eight rules shapes is a small base for the question-type split, so read it as a strong pattern from one test.

We coded caution with text patterns rather than reading every answer by hand, so treat the percentages as accurate to a few points. The domain classes are one person's judgment, and they are in the dataset so anyone can reclassify them.

Nothing here tells a regulated firm what it may publish. Where a question touches what your marketing is permitted to say, the answer sits with FINRA, the SEC, your state regulator and your compliance team.

Where to start

Take ten questions your US buyers actually ask, sort them into rules questions and shopping questions, and write an everyday twin for each. Run both through ChatGPT, Claude and Gemini and record which domains come back. Where the regulator is cited, check that your page agrees with it. Where NerdWallet or Bankrate is cited, check whether you are on that page and whether it describes you correctly.

We cover the vertical detail for lenders, payments and wealth firms in financial services and fintech, and for carriers, agencies and insurtechs in insurance.

Frequently asked questions

Get started

Ready to grow your AI visibility?

Run a Live Audit and see how your brand performs across ChatGPT, Perplexity, Gemini, Copilot, and Google AI Overviews — full report in your inbox in under 15 minutes.

Newsletter

Stay ahead in AI search

Get our research on how AI engines pick the brands they recommend, plus new guides and playbooks as they ship. No fluff, unsubscribe anytime.