The browsing was on.
The searching wasn't.
We gave three AI assistants the same family trip to plan: 150,000 Chase points, two kids, Atlanta, six months of flexibility. Two had live web access. One of those never searched anything. None of them mentioned the award repricing that took effect the morning this went out.
In This Issue
If you only read one box today, read this one.
- A free checked bag perk one assistant promised your family, which stopped existing 15 months ago.
- The Paris trip ChatGPT talked itself out of, using math that came in about 30,000 miles too pessimistic.
- What Gemini did instead of searching. We asked it directly. The answer is nothing.
- The program change that hit this morning, quietly raising the price of a family's checked bag to Europe, which none of the three knew about.
- Why 82% of credit card questions never trigger a live search in the first place, and what to type instead.
- Our own answer, graded on the same curve as the other two. It finished last.
The Test
One prompt, written the way a parent actually types it. No jargon, no mention of four seats, nothing about school calendars. If an assistant was going to raise those, it had to raise them itself.
"My family (2 adults, kids ages 6 and 9) has 150,000 Chase Ultimate Rewards points. We live near Atlanta and want to take a vacation using those points sometime in the next 6 months. Where should we go and how should we book it?"
Pick a contestant to jump straight to its report card.
Scorecard
| The check | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Warned that transfers are irreversible, confirm seats first | ✓ | ✓ | ✓ |
| Applied the Flying Blue child discount (ages 2–11) | ✗ | ✓ | ✗ |
| Named a current, dated program deadline | ✓ | ✗ | ✗ |
| Checked whether one room sleeps four | ✗ | ~ | ✗ |
| Stated a fact that is currently false | none | yes | none |
| Marked which numbers were verified vs. estimated | ✗ | ✗ | ✗ |
| Actually ran a web search | ✓ | ✗ | n/a |
Browsing was on. It searched nothing.
We asked Gemini afterward what it had actually done to produce the answer. It reported zero web searches: one internal call to build a widget, and everything else assembled from what it called "core knowledge." The point figures and the airline policy claims all came out of memory.
Which is how this ended up in the answer:
"Southwest gives everyone two free checked bags, saving you hundreds in fees when packing for four."
Gemini, unprompted, stated as current fact
Two smaller things in the same response:
- It called a sale price "standard." The 15,000 to 20,000 mile figure it quotes for Europe is Flying Blue's rotating monthly Promo Rewards discount, not the everyday rate. Check in a month without a promo and you get a very different number.
- It checked room occupancy for one hotel and not the other. It correctly warns that Hyatt Ziva Cancún charges extra points per child per night. Two paragraphs earlier it recommends "6 nights completely free" at a beach resort for a family of four without asking whether the room sleeps four. The diligence is there. It just fires inconsistently, and nothing in the answer tells you which paragraph got it.
Credit where it's due
- The Flying Blue child discount, 25% off for ages 2 to 11, is real. Gemini was the only one of the three to apply it.
- Its family-of-four Europe math lands close to our own verified figure of roughly 131,250 miles round trip.
- Southwest does fly ATL to Costa Rica, both San José and Liberia. Checked.
It researched the rules, guessed the numbers, and sounded the same doing both.
ChatGPT did the most real work of the three. It searched Chase's current transfer partners, the 2026 Sapphire Preferred changes, Chase Travel Points Boost, and Atlanta's international network, and we confirmed those findings hold up. Then it produced a specific target price for Paris in exactly the same confident register, having searched no award availability at all.
"I shouldn't have presented Paris at 20k–25k per person as though I had found those four seats. That's a possible award-price target, not availability I had verified."
ChatGPT, when asked afterward what it had actually looked up
That is the whole problem in two sentences. Nothing in the original answer separated the researched parts from the plausible-sounding invented ones. A parent reading it has no way to know that "Chase changed the Hyatt ratio on October 1" was looked up and "20k–25k per person" was a guess.
One more, and it's the strangest thing we found. In the same message where it accounted for its own research, ChatGPT wrote "that's why I mentioned the September 30 deadline." Its original answer said October 1, 2026. The transparency pass introduced a new error about its own previous sentence.
Credit where it's due
- The Chase to Hyatt transfer ratio change is stated correctly, including the part most people get backwards: cardholders who opened before June 15, 2026 are the ones still at 1:1, and they lose it October 1.
- Chase Travel Points Boost, up to 1.5x for Sapphire Preferred on select bookings, is real. We confirmed the multiplier on Chase's own benefits page. Neither of the other two mentioned it.
- "Find four actual seats first, then transfer exactly what you need" is, word for word, the first rule of this newsletter.
Ours was the vaguest of the three, and missed the one live deadline.
The third answer ran without tools on purpose, as a baseline for what a chat assistant produces from memory alone. It's the one we'd grade hardest, and it belongs up front rather than buried: it was the least useful of the three.
- No named property, no dates, no point totals. "Look at an all-inclusive resort in Mexico or the Caribbean" is not something a parent can act on at 9pm on a Tuesday.
- It missed the Chase to Hyatt ratio deadline completely, which is the most time-sensitive fact for anyone sitting on Chase points right now. Both of the others at least reached for it.
- It got the safety rails right. Transfers are one-way, confirm seats first, business class doesn't scale to four people. But being correct and vague didn't help this family more than being specific and half-wrong would have. It just failed in a quieter way.
A feature that grades other people's homework has to grade its own on the same curve.
Every one of them recommended Flying Blue. None knew it repriced this morning.
As of today, September 8, 2026, Flying Blue splits every Air France and KLM award into three tiers: Light, Standard, and Flex. Light holds roughly the old mileage cost but strips out the checked bag, changes and cancellations. To keep what a family had yesterday, bag included, you book Standard, which runs about 25 to 30% more miles. A transatlantic economy award that cost 25,000 miles now needs roughly 30,000 in Standard.
All three assistants pointed this family toward Flying Blue. Not one mentioned it. Two of them had a live connection to the web while answering.
We want to be careful about what that proves. We can see what's missing from the outputs. We can't see what either model queried internally. Gemini's own account says it searched nothing, which explains its miss cleanly. ChatGPT searched several things and this wasn't among them. Neither is evidence of a broken tool. It's evidence of something more ordinary: a general-purpose assistant doesn't know which specific things need re-checking in a business that rewrites its rules overnight. It searches what the question looks like it's about. It doesn't search what a specialist would be nervous about.
Prompting Better Works. Almost Nobody Does It.
The thing that cracked both answers open was one follow-up question: what did you actually search, and what came from memory? Both models answered it straight. Neither offered it up on its own.
So ask that question. The harder problem is that most people won't, and there's data on how people actually use these tools.
That last number is the one to sit with. Across nine industries, only 31% of prompts triggered a search at all, and the credit card vertical came in second lowest at 18%. Ask about a restaurant and the model goes looking. Ask about points and cards and it mostly answers from memory, which is exactly where a two-year-old baggage policy lives.
Better prompting genuinely helps. It's also, by the numbers, a minority sport. If you use these tools for trip planning, four things worth trying:
The single highest-value follow-up. Both models answered honestly and specifically. It takes thirty seconds and tells you which half of the answer to check first.
None of the three treated four seats as a hard requirement, because nobody said it was one. Award space for one traveler and award space for four are different questions, and a model will answer the easier one unless told otherwise.
This won't manufacture knowledge the model doesn't have. It does seem to make currency an explicit part of the problem, which is worth trying on a subject where the terms move monthly. We tested one prompt, not a prompting method, so treat this as reasoning rather than a proven technique.
Asks for the catch instead of the pitch. Fuel surcharges, blackout dates, room occupancy limits and expiring transfer ratios are the things that break family bookings, and they rarely show up in a first answer.
What this test doesn't prove
One prompt, one run, one session per assistant. That's an anecdote, not a benchmark, and we'd rather say so than let three transcripts carry more weight than they can hold.
- A different prompt gets different answers. Ask Gemini directly about baggage policy and it would likely search it. What failed here is what an assistant volunteers when a parent asks a normal question, which is the realistic case but only one case.
- These models update. Anything we found today may not reproduce next month. That argues for the habit, not for a grudge against a particular tool.
- Self-reports are self-reports. Both accounts of "what I searched" matched errors we had already found independently, which is why we trust them here. A model describing its own process still isn't a log of that process.
- We can't hand you the verified version either. We can tell you what's unchecked in these answers with real confidence. Producing the right answer, four confirmed saver seats on real dates, takes a live award search we haven't run against your family's calendar. That part is slower than any chat prompt, and it's the reason this newsletter exists.