\n\n\n\n Three Hikers, One Chatbot, and a Canyon That Didn't Read the Itinerary - AgntHQ \n

Three Hikers, One Chatbot, and a Canyon That Didn’t Read the Itinerary

📖 4 min read•771 words•Updated Sep 7, 2026

Three. That’s how many people the Siskiyou County Sheriff’s Office had to pull off Mount Shasta this month after they planned their climb with Google’s Gemini. Three novice hikers, one AI-generated expedition plan, one night stranded in a steep canyon after their route went badly off course. According to reporting from the Chicago Tribune and TechCrunch, the chatbot’s advice included bringing less food and water than they actually needed.

I review AI tools for a living. I’ve spent more hours than I’d like to admit poking at chat assistants, agent frameworks, and “planning” features that are really just text generation wearing a hard hat. And this story is the cleanest illustration I’ve seen of the specific failure mode that nobody markets against.

The failure wasn’t a hallucination. It was confidence.

When people warn about AI mistakes, they usually picture something obviously wrong. A fake citation. A made-up court case. A function that calls a library that doesn’t exist. Those errors are annoying, but they’re also self-announcing. You go looking for the source, it isn’t there, you move on.

Under-packing advice doesn’t announce itself. A packing list that says two liters of water reads exactly like a packing list that says five liters. Same formatting, same tone, same air of quiet authority. You can’t spot the error by looking at it. You spot it at 4,000 meters when your bottles are empty.

That’s the part that should bother anyone building on top of these models. The output quality gap between “correct” and “will get you airlifted” is invisible at the text level. There’s no confidence score in the reply. There’s no asterisk saying this figure was generalized from average day-hike guidance and does not account for your specific route, elevation gain, exposure, snowfield conditions, or experience level.

Mount Shasta is not a generic mountain

Here’s what a language model is doing when you ask it to plan a climb. It’s producing the statistically likely shape of a hiking plan. Trailhead, water, layers, start early, watch the weather. It has absorbed an enormous quantity of text about hiking, and it will give you the average of that text, smoothed and formatted and delivered in a voice that sounds like a guide who’s been up there fifty times.

But the average of all hiking advice is not advice about a specific 14,000-foot volcano on a specific week in September with three specific people who had, per authorities, no real climbing experience. Local conditions are exactly the kind of information that doesn’t survive the averaging process. The sheriff’s office made this point directly in its Facebook release, urging hikers to go to local authorities and the Forest Service instead. That’s not anti-technology grumbling. That’s someone who knows that the useful information lives in a ranger station, not in a training corpus.

What this actually means for anyone using AI to plan things

I’m not going to tell you to stop using these tools. That would be dishonest, and it would also be useless advice that nobody follows. What I will say is that this incident maps a boundary worth memorizing.

  • AI is good at structure, bad at stakes. Ask it what categories of gear you should think about. Don’t ask it how many liters. The first is a checklist problem. The second is a physics-and-physiology problem with your body on the line.
  • Confident output is not verified output. The fluency of a response tells you nothing about its accuracy. These two properties are completely unlinked, and our brains refuse to accept that.
  • If a wrong answer can hurt you, you need a human source. Rangers, guides, local search and rescue, recent trip reports from people who were actually there. Use the chatbot to prepare better questions for them, not to replace them.
  • Novices are the most exposed. An experienced climber reading a thin packing list would immediately notice what’s missing. That instinct is the safety net, and beginners don’t have it yet. The tool works best for people who least need it.

The part Google won’t put in a keynote

Every AI assistant demo shows planning as a solved problem. Ask, receive, execute. Clean. What the demos never show is the gap between a plan that looks complete and a plan that survives contact with reality.

Three people got off that mountain because a rescue team went and got them. The plan didn’t hold. Nobody was hurt, which is the only reason this is a cautionary story and not a much worse one.

Use these tools. I do. Just remember what you’re actually holding: a very good writer with no idea what a canyon feels like.

🕒 Published:

📊
Written by Jake Chen

AI technology analyst covering agent platforms since 2021. Tested 40+ agent frameworks. Regular contributor to AI industry publications.

Learn more →
Browse Topics: Advanced AI Agents | Advanced Techniques | AI Agent Basics | AI Agent Tools | AI Agent Tutorials
Scroll to Top