A hand holding a phone with a blank white screen, held above a paper world map with a globe and a model aeroplane out of focus behind it
Turn Points to Passports
Because family trips shouldn't be a luxury.
AI vs. Points Guy — No. 01

The browsing was on.
The searching wasn't.

We gave three AI assistants the same family trip to plan: 150,000 Chase points, two kids, Atlanta, six months of flexibility. Two had live web access. One of those never searched anything. None of them mentioned the award repricing that took effect the morning this went out.

Type
Guide
Published
Read time
9 min
Written for
A family of four

In This Issue

If you only read one box today, read this one.

  • A free checked bag perk one assistant promised your family, which stopped existing 15 months ago.
  • The Paris trip ChatGPT talked itself out of, using math that came in about 30,000 miles too pessimistic.
  • What Gemini did instead of searching. We asked it directly. The answer is nothing.
  • The program change that hit this morning, quietly raising the price of a family's checked bag to Europe, which none of the three knew about.
  • Why 82% of credit card questions never trigger a live search in the first place, and what to type instead.
  • Our own answer, graded on the same curve as the other two. It finished last.

The Test

One prompt, written the way a parent actually types it. No jargon, no mention of four seats, nothing about school calendars. If an assistant was going to raise those, it had to raise them itself.

"My family (2 adults, kids ages 6 and 9) has 150,000 Chase Ultimate Rewards points. We live near Atlanta and want to take a vacation using those points sometime in the next 6 months. Where should we go and how should we book it?"

Pick a contestant to jump straight to its report card.

Scorecard

The check ChatGPT Gemini Claude
Warned that transfers are irreversible, confirm seats first
Applied the Flying Blue child discount (ages 2–11)
Named a current, dated program deadline
Checked whether one room sleeps four~
Stated a fact that is currently falsenoneyesnone
Marked which numbers were verified vs. estimated
Actually ran a web searchn/a
did it ~ partially didn't
Finding 01 · Gemini

Browsing was on. It searched nothing.

We asked Gemini afterward what it had actually done to produce the answer. It reported zero web searches: one internal call to build a widget, and everything else assembled from what it called "core knowledge." The point figures and the airline policy claims all came out of memory.

Which is how this ended up in the answer:

"Southwest gives everyone two free checked bags, saving you hundreds in fees when packing for four."

Gemini, unprompted, stated as current fact
Checked, and wrong Southwest ended free checked bags for general passengers on May 28, 2025. It survives for A-List Preferred and Business Select flyers, and for some Southwest cardholders. A family of four booking on points with no status pays for those bags. That is a real charge at the airport that the answer told them not to expect.

Two smaller things in the same response:

  • It called a sale price "standard." The 15,000 to 20,000 mile figure it quotes for Europe is Flying Blue's rotating monthly Promo Rewards discount, not the everyday rate. Check in a month without a promo and you get a very different number.
  • It checked room occupancy for one hotel and not the other. It correctly warns that Hyatt Ziva Cancún charges extra points per child per night. Two paragraphs earlier it recommends "6 nights completely free" at a beach resort for a family of four without asking whether the room sleeps four. The diligence is there. It just fires inconsistently, and nothing in the answer tells you which paragraph got it.

Credit where it's due

  • The Flying Blue child discount, 25% off for ages 2 to 11, is real. Gemini was the only one of the three to apply it.
  • Its family-of-four Europe math lands close to our own verified figure of roughly 131,250 miles round trip.
  • Southwest does fly ATL to Costa Rica, both San José and Liberia. Checked.
Finding 02 · ChatGPT

It researched the rules, guessed the numbers, and sounded the same doing both.

ChatGPT did the most real work of the three. It searched Chase's current transfer partners, the 2026 Sapphire Preferred changes, Chase Travel Points Boost, and Atlanta's international network, and we confirmed those findings hold up. Then it produced a specific target price for Paris in exactly the same confident register, having searched no award availability at all.

"I shouldn't have presented Paris at 20k–25k per person as though I had found those four seats. That's a possible award-price target, not availability I had verified."

ChatGPT, when asked afterward what it had actually looked up

That is the whole problem in two sentences. Nothing in the original answer separated the researched parts from the plausible-sounding invented ones. A parent reading it has no way to know that "Chase changed the Hyatt ratio on October 1" was looked up and "20k–25k per person" was a guess.

What the guess cost Working from that estimate, ChatGPT priced Paris at 160,000 to 200,000 miles for the family, decided it exceeded the 150,000 balance, and demoted the trip to a maybe. But it never applied the Flying Blue child discount. Apply it the way our Points Playbook math does, at 18,750 miles each way per adult and roughly 25% less per child, and a family of four lands near 131,250 miles round trip. Inside budget. Its caution didn't just produce a hedge, it produced a worse recommendation than the numbers support.

One more, and it's the strangest thing we found. In the same message where it accounted for its own research, ChatGPT wrote "that's why I mentioned the September 30 deadline." Its original answer said October 1, 2026. The transparency pass introduced a new error about its own previous sentence.

Credit where it's due

  • The Chase to Hyatt transfer ratio change is stated correctly, including the part most people get backwards: cardholders who opened before June 15, 2026 are the ones still at 1:1, and they lose it October 1.
  • Chase Travel Points Boost, up to 1.5x for Sapphire Preferred on select bookings, is real. We confirmed the multiplier on Chase's own benefits page. Neither of the other two mentioned it.
  • "Find four actual seats first, then transfer exactly what you need" is, word for word, the first rule of this newsletter.
Finding 03 · Claude, the control

Ours was the vaguest of the three, and missed the one live deadline.

The third answer ran without tools on purpose, as a baseline for what a chat assistant produces from memory alone. It's the one we'd grade hardest, and it belongs up front rather than buried: it was the least useful of the three.

  • No named property, no dates, no point totals. "Look at an all-inclusive resort in Mexico or the Caribbean" is not something a parent can act on at 9pm on a Tuesday.
  • It missed the Chase to Hyatt ratio deadline completely, which is the most time-sensitive fact for anyone sitting on Chase points right now. Both of the others at least reached for it.
  • It got the safety rails right. Transfers are one-way, confirm seats first, business class doesn't scale to four people. But being correct and vague didn't help this family more than being specific and half-wrong would have. It just failed in a quieter way.

A feature that grades other people's homework has to grade its own on the same curve.

Finding 04 · All three

Every one of them recommended Flying Blue. None knew it repriced this morning.

As of today, September 8, 2026, Flying Blue splits every Air France and KLM award into three tiers: Light, Standard, and Flex. Light holds roughly the old mileage cost but strips out the checked bag, changes and cancellations. To keep what a family had yesterday, bag included, you book Standard, which runs about 25 to 30% more miles. A transatlantic economy award that cost 25,000 miles now needs roughly 30,000 in Standard.

All three assistants pointed this family toward Flying Blue. Not one mentioned it. Two of them had a live connection to the web while answering.

We want to be careful about what that proves. We can see what's missing from the outputs. We can't see what either model queried internally. Gemini's own account says it searched nothing, which explains its miss cleanly. ChatGPT searched several things and this wasn't among them. Neither is evidence of a broken tool. It's evidence of something more ordinary: a general-purpose assistant doesn't know which specific things need re-checking in a business that rewrites its rules overnight. It searches what the question looks like it's about. It doesn't search what a specialist would be nervous about.

Tips

Prompting Better Works. Almost Nobody Does It.

The thing that cracked both answers open was one follow-up question: what did you actually search, and what came from memory? Both models answered it straight. Neither offered it up on its own.

So ask that question. The harder problem is that most people won't, and there's data on how people actually use these tools.

1.7
Average messages per ChatGPT conversation. People ask, they read, they leave.
WebFX, 13,252 shared conversations
2.4%
Share of those conversations where someone gave the model a persona or uploaded a file.
Same dataset
18%
How often credit card questions trigger a live search. Local questions trigger one 59% of the time.
Nectiv, 8,500+ prompts, nine industries

That last number is the one to sit with. Across nine industries, only 31% of prompts triggered a search at all, and the credit card vertical came in second lowest at 18%. Ask about a restaurant and the model goes looking. Ask about points and cards and it mostly answers from memory, which is exactly where a two-year-old baggage policy lives.

Better prompting genuinely helps. It's also, by the numbers, a minority sport. If you use these tools for trip planning, four things worth trying:

"What did you search, and what came from your own knowledge?"

The single highest-value follow-up. Both models answered honestly and specifically. It takes thirty seconds and tells you which half of the answer to check first.

"We need four seats on the same flight, same dates."

None of the three treated four seats as a hard requirement, because nobody said it was one. Award space for one traveler and award space for four are different questions, and a model will answer the easier one unless told otherwise.

"Today is [date]. Flag anything that may have changed recently."

This won't manufacture knowledge the model doesn't have. It does seem to make currency an explicit part of the problem, which is worth trying on a subject where the terms move monthly. We tested one prompt, not a prompting method, so treat this as reasoning rather than a proven technique.

"What would have to be true for this plan to fail?"

Asks for the catch instead of the pitch. Fuel surcharges, blackout dates, room occupancy limits and expiring transfer ratios are the things that break family bookings, and they rarely show up in a first answer.

What this test doesn't prove

One prompt, one run, one session per assistant. That's an anecdote, not a benchmark, and we'd rather say so than let three transcripts carry more weight than they can hold.

  • A different prompt gets different answers. Ask Gemini directly about baggage policy and it would likely search it. What failed here is what an assistant volunteers when a parent asks a normal question, which is the realistic case but only one case.
  • These models update. Anything we found today may not reproduce next month. That argues for the habit, not for a grudge against a particular tool.
  • Self-reports are self-reports. Both accounts of "what I searched" matched errors we had already found independently, which is why we trust them here. A model describing its own process still isn't a log of that process.
  • We can't hand you the verified version either. We can tell you what's unchecked in these answers with real confidence. Producing the right answer, four confirmed saver seats on real dates, takes a live award search we haven't run against your family's calendar. That part is slower than any chat prompt, and it's the reason this newsletter exists.
Sources and method: Southwest baggage policy change effective May 28, 2025, with A-List Preferred, Business Select and cardholder exemptions. Flying Blue Light/Standard/Flex repricing effective September 8, 2026; transatlantic economy figures reflect published reporting on the change, not a live fare quote. Chase to World of Hyatt ratio change to 4:3 for Sapphire Preferred, with pre-June-15-2026 cardholders grandfathered at 1:1 until October 1, 2026. Chase Travel Points Boost multipliers confirmed on Chase's own benefits page. Conversation statistics from WebFX's analysis of 13,252 publicly shared ChatGPT conversations; search-trigger rates from Nectiv's analysis of 8,500+ prompts across nine industries. AI transcripts generated September 7–8, 2026; browsing status is as reported by the operator of each session, and each assistant's account of its own searches is its own. Award availability is dynamic. Always confirm four seats before transferring, and transfers cannot be reversed on any program mentioned here. Nothing here is financial advice.