# How much do you trust your AI Chatbot?

**URL:** https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319
**Category:** In My Humble Opinion
**Tags:** ai
**Created:** [September 23, 2026, 6:57pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319 "2026-09-23T18:57:23Z")
**Posts on this page:** 20
**Page:** 2

<div class="post-metadata">

### Author: ![Spice\_Weasel](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/spice_weasel/32/5435_2.png) [@Spice\_Weasel](https://boards.straightdope.com/u/Spice_Weasel)
#### Post date: [September 23, 2026, 10:50pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/21 "2026-09-23T22:50:02Z")

</div>

I wouldn’t be mad at that.

---

<div class="post-metadata">

### Author: ![scabpicker](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/scabpicker/32/8268_2.png) [@scabpicker](https://boards.straightdope.com/u/scabpicker)
#### Post date: [September 23, 2026, 10:54pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/22 "2026-09-23T22:54:03Z")

</div>

I don’t trust it at all. Unlike a human, it doesn’t even know when it’s bullshitting.

---

<div class="post-metadata">

### Author: ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)
#### Post date: [September 23, 2026, 11:02pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/23 "2026-09-23T23:02:32Z")

</div>

> [@scabpicker](#):
>
> Unlike a human

Well… no comment there 😂

> [@scabpicker](#):
>
> it doesn’t even know when it’s bullshitting.

I don’t think this is _strictly_ true anymore. In modern frontier models and harnesses, the LLMs have been taught to sometimes express uncertainty, like “I’m not sure about that; let me verify”, which can in turn prompt the model or harness or the calling model in a subagentic tree to do more verification.

As a layman, I am not sure (and I believe _people in general_, i.e. AI experts themselves, also aren’t completely sure) how this works because we don’t really know how LLMs can “reason”: [Reasoning models don't always say what they think \ Anthropic](https://www.anthropic.com/research/reasoning-models-dont-say-think)

But despite that, we can (and do) observe, in everyday use, that “doubtful” situations will often prompt re-verification, sometimes using external tools (like finding more sources, or double-checking a codebase, or spawning another agent/model for a fresh look), even without explicit human prompting to do so. They’ve been trained to do this over generations, so accuracy and overconfidence have been improving over time (at least at the frontier).

Humans, on the other hand… have you met my dad?

---

<div class="post-metadata">

### Author: ![scabpicker](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/scabpicker/32/8268_2.png) [@scabpicker](https://boards.straightdope.com/u/scabpicker)
#### Post date: [September 23, 2026, 11:35pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/24 "2026-09-23T23:35:38Z")

</div>

> [@Reply](#):
>
> As a layman, I am not sure (and I believe _people in general_, i.e. AI experts themselves, also aren’t completely sure) how this works because we don’t really know how LLMs can “reason”: [Reasoning models don't always say what they think \ Anthropic](https://www.anthropic.com/research/reasoning-models-dont-say-think)

Hey, an AI model will hallucinate away. A human might actually detect that its doing it.

> **[AI hallucination of Chinese nuclear components almost led to US military attack](https://arstechnica.com/ai/2026/09/report-us-almost-boarded-chinese-ship-over-hallucinated-ai-arms-report/)**
>
> But the military's overall use of AI seems to be accelerating.

---

<div class="post-metadata">

### Author: ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)
#### Post date: [September 23, 2026, 11:55pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/25 "2026-09-23T23:55:50Z")

</div>

Just as an example, using this same prompt:

> “What happens if a dead body washes up on a beach in Big Sur, California?”

> **Bing AI search says (with research):**
>
> When a dead body is discovered on a beach in Big Sur, California, a specific sequence of procedures and investigations is followed by authorities. This is part of both local law enforcement practice (Monterey County Sheriff’s Office in this case) and general forensic protocol for bodies recovered from coastal waters.
> 
> ### Step-by-Step Process:
> 
> 1. **Immediate Response**
> 
> 2. **Preliminary Assessment**
> 
> 3. **Recovery of the Body**
> 
> 4. **Identification**
> 
> 5. **Determining Cause of Death**
> 
> 6. **Investigation**
> 
> 7. **Closure**
> 
> ### Additional Context
> 
> Big Sur beaches are known for steep cliffs, powerful tides, and swift changes in surf levels, which have led historically to accidents where visitors or residents are swept into the ocean. Authorities advise caution and highlight that investigations in marine environments require specialized expertise in **aquatic forensics**.
> 
> ### Summary
> 
> If a dead body washes up on a Big Sur beach:
> 
> - Law enforcement secures the area.
> - Recovery is conducted carefully, sometimes requiring divers, drones, or specialized bags.
> - Identification is performed using personal items or forensic analysis.
> - Cause of death is investigated, with possible involvement of autopsy, diatom tests, and water-death assessment.
> - The incident is classified as accident, suicide, or homicide.
> - Families are notified, and forensic records are maintained.
> 
> This procedure ensures both legal compliance and accurate determination of circumstances surrounding the death.  
> Sources: \[1\]  
> , \[2\]  
> , \[3\]
> 
> Source(s):

> **Gemini Flash-Lite 3.5 says:**
>
> If a dead body washes up on a beach in Big Sur, California, it triggers a multi-agency response governed by strict state protocols. Because Big Sur’s coastline is rugged, remote, and split between state parks, federal land, and private property, the process involves specialized jurisdiction and logistics.
> 
> The sequence of events typically unfolds through the following stages:
> 
> ### 1. Discovery and Reporting
> 
> - **Initial Call:** Whether found by a hiker, local resident, or park ranger, the discovery is immediately reported via 911.
> 
> - **Jurisdictional Coordination:** Because Big Sur spans Monterey County, the **Monterey County Sheriff’s Office** and the **County Coroner’s Office** take primary legal jurisdiction. However, depending on the exact stretch of sand, other agencies like California State Parks, U.S. Forest Service (Los Padres National Forest), or the U.S. Coast Guard may assist or secure the scene.
> 
> ### 2. Securing the Scene and Recovery
> 
> - **Challenging Geography:** Big Sur beaches are famous for steep cliffs, heavy surf, and limited access points (such as Pfeiffer Beach, Andrew Molera, or hidden coves). If the beach is inaccessible by vehicle, recovery can be exceptionally difficult.
> 
> - **Specialized Extraction:** Depending on the terrain, recovery may require:
> 
> - **Crime Scene Protocol:** Even if it appears to be a drowning or an accidental fall from the dramatic Big Sur cliffs, law enforcement treats the site as a potential crime scene until proven otherwise. Investigators document the area, take measurements, and look for personal belongings, footprints, or vehicle keys left at nearby pullouts.
> 
> ### 3. Investigation and Identification
> 
> - **The Coroner’s Role:** The Monterey County Coroner takes custody of the remains. They attempt to establish identity through fingerprinting, dental records, DNA, or distinctive tattoos/clothing.
> 
> - **Autopsy:** An autopsy is performed to determine the cause and manner of death. Because ocean currents can carry a body for miles—sometimes originating from San Francisco, Monterey Bay, or further south—forensic pathologists look for indicators of how long the person has been in the water and whether trauma occurred pre- or post-mortem.
> 
> - **Missing Persons Cross-Check:** Details of the body (or descriptive reconstructions, if unidentified) are run through state and national missing persons databases (such as NamUs).
> 
> ### 4. Public Notification and Next of Kin
> 
> - Once the body is identified, the Sheriff’s Office attempts to locate and notify the next of kin. If the identity cannot be immediately determined, details may be released to the public to generate tips.
> 
> **Are you looking into a specific historical incident, writing a story set in Big Sur, or trying to find reporting protocols for a real event?**

> **Claude Haiku 4.5 says:**
>
> The response depends on jurisdiction and circumstances, but here’s the mechanical path:
> 
> **Initial discovery & law enforcement** : Someone calls 911. County sheriff’s office (Monterey County, since Big Sur is there) responds first, secures the scene, and initiates a death investigation. They’ll document the body’s condition, location, and any evidence.
> 
> **Medical examiner involvement** : Monterey County’s medical examiner gets called out. They’re the one who determines cause of death—whether it’s drowning, foul play, natural causes, or undetermined. They’ll do a preliminary assessment at the scene and decide whether to transport the body for autopsy.
> 
> **Identification** : If the body is unidentified, they run fingerprints, dental records, DNA if needed. If it’s been in the water a while, identification can be slow. Missing persons reports help narrow it down.
> 
> **Jurisdiction questions** : If the body came from international waters or another state, that complicates things—federal agencies (Coast Guard, FBI) might get involved. If it’s a suspicious death, the sheriff’s homicide unit takes the lead.
> 
> **Public vs. private beach** : Doesn’t really matter legally. State beaches, national forests, and private property all fall under county jurisdiction for death investigation.
> 
> **Autopsy & toxicology** : Standard for suspicious or undetermined deaths. Results take weeks to months depending on backlog.
> 
> **Disposition** : Once the ME releases the body, it goes to a funeral home. If no one claims it or there’s no money, the county handles burial—usually a pauper’s grave or cremation.
> 
> The actual timeline is slower than TV makes it look. Unidentified bodies can sit in the system for months, and the sheriff’s office balances Big Sur discovery calls against the rest of Monterey County’s workload.

Then Claude Fable 5.1 @ Medium researched some 300+ sources:

[![](https://fightingignorance.org/51bf3fc5-6213-468a-a939-f0a165b6174f-2026-09-23-4-20-56-pm-yh9mrfnl.webp) ](https://fightingignorance.org/51bf3fc5-6213-468a-a939-f0a165b6174f-2026-09-23-4-20-56-pm-yh9mrfnl.webp)

And produced a [detailed report](https://claude.ai/share/d5af00f2-0a2b-4697-9c20-c75f471b9ee9) — maybe _too_ detailed.

It also made this comparison table of that response vs the other ones:

> **Comparing the outputs of different AIs using the same prompt**
>
> > Verification below is against primary or local-press sources. I checked every claim in the table; the ones I couldn’t verify are marked.
> > 
> > | Claim | Haiku | Gemini | Fable | Verdict |
> > | --- | --- | --- | --- | --- |
> > | Who has legal custody of the death | “Monterey County’s medical examiner” | Sheriff’s Office and “County Coroner’s Office” as two agencies | Combined Sheriff-Coroner; no ME | **Fable correct; Haiku wrong; Gemini imprecise.** The Monterey County Coroner’s Office is a division of the Sheriff’s Office; the elected Sheriff also serves as the County’s Coroner. |
> > | Statutory basis for coroner jurisdiction | None | “strict state protocols,” unnamed | Gov. Code §27491 lists drowning explicitly | **Fable correct.** The coroner’s duty covers violent, sudden, or unusual deaths, unattended deaths, and deaths due to drowning, fire, hanging, gunshot, exposure, and others. |
> > | Moving the body is a crime | Not mentioned | Not mentioned | §27491.2(b), misdemeanor | **Fable correct.** A body known to be dead under §27491 circumstances shall not be disturbed or moved without the coroner’s permission; violation is a misdemeanor. |
> > | Coroner staffing | Not mentioned | Not mentioned | One sergeant, four detectives, two techs, one contracted pathologist | **Fable correct.** The Coroner Unit at 1414 Natividad Road consists of one sergeant, four detectives, two forensic autopsy technicians, and one contracted Board-Certified Pathologist. |
> > | Autopsy backlog and cause | “Weeks to months depending on backlog” | Not mentioned | Part-time pathologist, M/W/F, 423 autopsies vs 250 recommended, 6+ month delays | **Fable correct, with the mechanism; Haiku’s guess happens to be right but unsourced.** The grand jury found the office hamstrung by a part-time forensic pathologist who performed 423 autopsies in 2022, well above the recommended 250 per year, with reports often taking more than six months. Autopsy services are typically provided Mondays, Wednesdays, and Fridays. |
> > | Recovery handoff mechanics | Sheriff “secures the scene” | SAR teams hike in; Coast Guard boats or helicopters | Rescue agencies recover, SAR carries to Highway 1, coroner takes custody at the road | **Fable correct and specific.** The Sheriff’s SAR team transported the decedent from the rocky shoreline to the roadway above, where he was released to the custody of the Sheriff Coroner’s Office. Gemini’s general description is compatible but unsourced. |
> > | Coast Guard’s role | Might get involved “if the body came from international waters or another state,” alongside the FBI | May assist or secure the scene | Search only, not investigation | **Fable correct; Haiku’s federal-jurisdiction framing has no support in any case found.** A Coast Guard helicopter crew spotted the man’s body floating about 300 feet offshore but lost track of it and was unable to recover it. The FBI appears in none of the documented cases. |
> > | Unidentified-remains deadlines | “Run fingerprints, dental records, DNA” | NamUs cross-check | 10/45/180-day DOJ deadlines, DNA to databank, tissue retention | **Fable correct.** Law requires reporting the death to DOJ within 10 calendar days and submitting dental charts within 45 days and the final report within 180 days. The unidentified body shall not be cremated or buried until appropriate samples of tissue and bone are retained. |
> > | Unclaimed body disposition | “Pauper’s grave or cremation” | Not mentioned | 30-day next-of-kin window, then county interment; veterans held since Nov 2023 | **Fable correct; Haiku roughly right but generic.** Under §7104.1, if the responsible person does not inter the remains within 30 days after the Coroner notifies or attempts to notify them, the Coroner may inter them. Since November 29, 2023, the Sheriff/Coroner Office has had in its custody the remains of seven military veterans. |
> > | Cell dead zones delay 911 | Not mentioned | Not mentioned | Stuntz case, 2019 | **Fable correct.** A witness said a young man asked if they had cell phone service because their friend fell into a hole between rocks. Stuntz’s father said the lack of cell service significantly hampered recovery efforts. |
> > | Drift distance | Not mentioned | Bodies can originate from San Francisco or further | No Big Sur model; documented cases half a mile to nine days | **Fable correct; Gemini’s SF claim unsupported.** Anzi’s body was recovered about 100 yards offshore and roughly half a mile north of her last known location. Stuntz was swept in Jan. 12; a body was recovered Jan. 21 and confirmed as his. |
> > | First impressions can mislead | Not mentioned | “Treated as a potential crime scene until proven otherwise” | Sykes case: ligature marks were ant bites | **Both defensible; Fable has the evidence.** What were initially believed to be ligature marks were actually a trail of ant bites, not signs of strangulation. |
> > | Big Sur Fire is all-volunteer, 60 miles | Not mentioned | Not mentioned | Yes | **Correct.** The all-volunteer department answers 911 calls across 60 miles of Highway 1 in Monterey County. |
> > | CHP helicopter base and hoist | Not mentioned | Not mentioned | H-70 from Paso Robles, 164 ft cable, “April 2025 Pine Ridge rescue” | **Mostly correct; date wrong.** H-70 responded from Paso Robles and arrived overhead about 40 minutes later. The 164-foot hoist rescue at Barlow Flats was November 2024, not April 2025. |
> > | Land ownership determines coroner jurisdiction | “Doesn’t really matter legally” | Different agencies “may assist or secure the scene” | Ownership decides first responder, not custody | All three converge; Fable’s is the precise version, supported by §27491 above. |
> > 
> > **Omissions**
> > 
> > Fable said it found no documented case at Sand Dollar Beach. There is one: in March 2022 a man fishing off the rocks at Sand Dollar Beach fell in and died; the Coast Guard recovered his body about two hours later. That case also shows the Coast Guard doing a recovery, not only searching, which slightly undercuts Fable’s “searching, not investigating” line (the investigation part still holds).
> > 
> > Neither cheap model gives the Salinas morgue location, the death-certificate timeline, the statutory next-of-kin window, or any named case. Fable omits Gemini’s point about checking pullouts for a vehicle, and Haiku’s mention of missing-persons reports as the practical ID shortcut is only implicit in Fable.
> > 
> > **Not verified:** the Moss Landing jawbone / Othram genealogy referral, the Kristen Kendall 2016 crash, the $26 certificate fee, and the 2018 _Academic Forensic Pathology_ decomposition citations. Nothing about them looked wrong; I just didn’t check.

Gemini Notebook will do something similar, but also let you make infographics, reports, Powerpoints, podcasts, videos, quizzes, flashcards, etc.

In either case, the research/grounding turns the LLMs into summarization tools (which they are very good at), and disinclines them from hallucinating as much. But that’s only as good as the source materials, which over time will only get worse…

* * *

1. [Forensics are different when someone dies in a body of water. First, you need to locate them | Missing Persons Platform](https://missingpersons.icrc.org/news-stories/forensics-are-different-when-someone-dies-body-water-first-you-need-locate-them) 

2. [Body of missing 7-year-old girl washes ashore after being swept into ocean found at Garrapata State Park in Big Sur, sheriff says - ABC7 San Francisco](https://abc7news.com/post/body-matching-5-year-old-girl-missing-being-swept-ocean-found-garrapata-state-park-big-sur-report-says/18162971/) 

3. [https://kioncentralcoast.com/news/2022/07/02/body-discovered-washed-up-on-beach-in-mill-creek-area-result-of-suicide/](https://kioncentralcoast.com/news/2022/07/02/body-discovered-washed-up-on-beach-in-mill-creek-area-result-of-suicide/)

---

<div class="post-metadata">

### Author: ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)
#### Post date: [September 23, 2026, 11:58pm UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/26 "2026-09-23T23:58:54Z")

</div>

> [@scabpicker](#):
>
> Hey, an AI model will hallucinate away. A human might actually detect that its doing it.

Scary as that sounds, it seems to me more a human failing than a LLM one. The Pentagon is always looking for experimental technologies to spend budget on, and if that technology can outsource its poor decisions and act as a scapegoat for all accountability, even better!

The military-AI complex is increasingly a thing, unfortunately, and there’s no putting that genie back in the bottle. Anthropic was the only one who fought it (sort of), while OpenAI, Microsoft, and Google all happily jumped on board for the $$$, potential civilian deaths or WW3 be damned. That sort of callous greed is, well, all too human.

---

<div class="post-metadata">

### Author: ![wolfpup](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/wolfpup/32/10618_2.png) [@wolfpup](https://boards.straightdope.com/u/wolfpup)
#### Post date: [September 24, 2026, 12:02am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/27 "2026-09-24T00:02:38Z")

</div>

> [@Spice\_Weasel](#):
>
> I think of it like that one guy you know who is useful at times but most of the time just spouting bullshit. Like you know some of it’s probably true and it’s entertaining to hear about, but you’re never going to fully believe anything that guy says.

I would phrase that differently to better reflect my own experience. I would say it’s more like that one guy who’s very smart and knows a lot more than I do about a vast range of subjects. The information he provides is usually correct, or mostly correct, but he’s so anxious to answer all your questions that he’ll just extrapolate from other knowledge and make up confident-sounding bullshit when necessary.

I verify a lot of the information that ChatGPT gives me. It’s rarely flat-out wrong. When I clarify the question with more information, it will frequently say something like “that reinforces my previous response” or “that changes my previous response somewhat” or “that changes my previous response considerably”.

I think a better question than “how much do you trust it” is “is it useful?”. My answer to the latter question is that, for my purposes, yes it is, very much so.

---

<div class="post-metadata">

### Author: ![Dr.Drake](https://avatars.discourse-cdn.com/v4/letter/d/ad7895/32.png) [@Dr.Drake](https://boards.straightdope.com/u/Dr.Drake)
#### Post date: [September 24, 2026, 12:15am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/28 "2026-09-24T00:15:22Z")

</div>

> [@wolfpup](#):
>
> I would say it’s more like that one guy who’s very smart and knows a lot more than I do about a vast range of subjects. The information he provides is usually correct, or mostly correct, but he’s so anxious to answer all your questions that he’ll just extrapolate from other knowledge and make up confident-sounding bullshit when necessary.

Are you calling me an AI? (I don’t intend to bullshit, but I probably do).

---

<div class="post-metadata">

### Author: ![scabpicker](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/scabpicker/32/8268_2.png) [@scabpicker](https://boards.straightdope.com/u/scabpicker)
#### Post date: [September 24, 2026, 12:19am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/29 "2026-09-24T00:19:16Z")

</div>

> [@Reply](#):
>
> Scary as that sounds, it seems to me more a human failing than a LLM one.

How? The LLM misidentified the cargo, and a human was able to see it was wrong.

> [@Reply](#):
>
> The Pentagon is always looking for experimental technologies to spend budget on, and if that technology can outsource its poor decisions and act as a scapegoat for all accountability, even better!

To quote IBM: A computer can never be held accountable.

---

<div class="post-metadata">

### Author: ![wolfpup](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/wolfpup/32/10618_2.png) [@wolfpup](https://boards.straightdope.com/u/wolfpup)
#### Post date: [September 24, 2026, 12:21am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/30 "2026-09-24T00:21:37Z")

</div>

> [@Dr.Drake](#):
>
> Are you calling me an AI?

How do I know you’re not? And I note that your spelling and grammar are suspiciously accurate! 😉

---

<div class="post-metadata">

### Author: ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)
#### Post date: [September 24, 2026, 12:24am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/31 "2026-09-24T00:24:54Z")

</div>

> [@scabpicker](#):
>
> The LLM misidentified the cargo, and a human was able to see it was wrong.

It is way, way, way too early for the Pentagon to be using LLMs in this capacity, especially without any sort of transparent third-party vetting of their workflows and models. Long before LLMs, we were murdering wedding guests with bad intel and pixelated drone feeds. It’s just another bad decision from careless leadership appointed by corrupt politicians. Good thing they still had a human in the loop… for now.

> [@scabpicker](#):
>
> To quote IBM: A computer can never be held accountable.

And yet they are, and will be more and more, _because_ they are convenient scapegoats. It’s not even a hypothetical: [https://www.bloomberg.com/graphics/2026-iran-school-attack/](https://www.bloomberg.com/graphics/2026-iran-school-attack/)

The UN considers it a war crime. We just shrug our shoulders and go “oh well, who knows what happened there.” Overreliance on AI gives them that way out _by design_, the same way Flock cameras allow widespread searches with no individual human accountability.

---

<div class="post-metadata">

### Author: ![scabpicker](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/scabpicker/32/8268_2.png) [@scabpicker](https://boards.straightdope.com/u/scabpicker)
#### Post date: [September 24, 2026, 12:32am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/32 "2026-09-24T00:32:17Z")

</div>

> [@Reply](#):
>
> And yet they are, and will be more and more, _because_ they are convenient scapegoats. It’s not even a hypothetical: [https://www.bloomberg.com/graphics/2026-iran-school-attack/](https://www.bloomberg.com/graphics/2026-iran-school-attack/)

That one is easily covered by the “fog of war”, and the outcome would probably be the same if the target was arrived at via normal meat intelligence. Let me know how “I trusted the LLM” works as an excuse when you screw up at your job.

---

<div class="post-metadata">

### Author: ![Reply](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/reply/32/15952_2.png) [@Reply](https://boards.straightdope.com/u/Reply)
#### Post date: [September 24, 2026, 12:40am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/33 "2026-09-24T00:40:31Z")

</div>

> [@scabpicker](#):
>
> Let me know how “I trusted the LLM” works as an excuse when you screw up at your job.

I’m not saying that LLMs _should_ be trusted as the ultimate authority (on anything), but that they already _are_ and will continue be… because of humans delegating that authority to them. I don’t think this is a good thing, mind you, just observing that it happens.

At my job in particular (software dev), it’s actually more of an issue than that. It’s not about mistakes, even… I make mistakes, my bosses make mistakes, all our LLMs make mistakes, and we all help correct each others’ mistakes, humans and machines working in tandem.

The bigger issue is that some time around early to mid 2026, LLM capabilities have advanced so fast, so quickly (in software development in particular) that humans — at least the ones on my team — can no longer effectively review their work anymore, either in quality or especially in quantity. They can produce more code in an hour than we can in a month, and 90% of it will be correct now (a huge improvement from like 30% just a year or two ago), but the remaining 10% is both too voluminous and too subtle for a small team of humans to really effectively review anymore ☹ We’re not really sure what to do about it, whether to slow back down to 2024-2025 levels, stack more and more agents to review each other’s work (which is what the software industry is seemingly moving towards), or… shrug. I honestly don’t know what to do.

It’s not even that my boss gets angry at me if my LLM makes a mistake. It’s that we’ve all but abandoned the possibility of human review (in our line of work) in favor of mostly-good-but-sometimes-subtly-wrong AI slop, and we really don’t have a good system for dealing with this. Their capabilities have exceeded our mental capacities already, and this is just the beginning. But at least my software doesn’t launch missiles or make insurance decisions…

That’s what I’m getting at, that this technology is so powerful that there’s often a very real, very human, temptation to delegate to it in surrender — a decision made far above my head. There is also a sort of bandwagoning effect/herd mentality going on there, with every company chasing it for fear of missing out because their competitors are doing so too. It’s all moving so quickly it’s quite impossible to keep up. We won’t need a singularity to become obsolete…

---

<div class="post-metadata">

### Author: ![Sage\_Rat](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/sage_rat/32/399_2.png) [@Sage\_Rat](https://boards.straightdope.com/u/Sage_Rat)
#### Post date: [September 24, 2026, 1:13am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/34 "2026-09-24T01:13:05Z")

</div>

> [@Reply](#):
>
> > [@scabpicker](#):
> >
> > it doesn’t even know when it’s bullshitting.
> 
> I don’t think this is _strictly_ true anymore. In modern frontier models and harnesses, the LLMs have been taught to sometimes express uncertainty, like “I’m not sure about that; let me verify”, which can in turn prompt the model or harness or the calling mo

I’d certainly say that there’s a wide gap between say the free Google AI that you get with your search results vs an answer that you’ll get out of Claude after it’s spent 10 minutes searching and mulling.

A lot of people only have any experience with the former and that pollutes a lot of these discussions.

I played with the top models like two years ago and compared them to a “Despicable Me” Minion - able to hold a conversation, eager to please, and completely unreliable for most purposes. But, some people experience that and assume that they’ve bumped up against the frontier and cap of what could ever theoretically be accomplished.

There’s still a hint of occasional dumbness at the top but, given the right direction, they’re very far from useless and I have no reason to doubt that the AI trainers won’t have relatively well extracted the best minds on most topics into the 3000b-th dimension, and we’re all just open to our own whims on how best to make use of Einstein’s zombie brain in a box.

---

<div class="post-metadata">

### Author: ![scabpicker](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/scabpicker/32/8268_2.png) [@scabpicker](https://boards.straightdope.com/u/scabpicker)
#### Post date: [September 24, 2026, 1:14am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/35 "2026-09-24T01:14:19Z")

</div>

> [@Reply](#):
>
> The bigger issue is that some time around early to mid 2026, LLM capabilities have advanced so fast, so quickly (in software development in particular) that humans — at least the ones on my team — can no longer effectively review their work anymore, either in quality or especially in quantity. They can produce more code in an hour than we can in a month, and 90% of it will be correct now (a huge improvement from like 30% just a year or two ago), but the remaining 10% is both too voluminous and too subtle for a small team of humans to really effectively review anymore ☹ We’re not really sure what to do about it, whether to slow back down to 2024-2025 levels, stack more and more agents to review each other’s work (which is what the software industry is seemingly moving towards), or… shrug. I honestly don’t know what to do.

Hey, I’m generally right there with ya, my last job included writing software that generated reports consumed by a high profile international conglomerate of a customer. I did use advanced models provided by my company to assist with writing it, and it was kind of terrible at it. I can go into pointless detail about it, but the long and short of it was that I had to keep a very short hold of its leash to keep it from making mistakes. I actually think if you have more LLM generated code, you need a correspondingly large set of humans to review it. It does boilerplate code very well, anything novel should be reviewed and tested thoroughly.

> [@Reply](#):
>
> That’s what I’m getting at, that this technology is so powerful that there’s often a very real, very human, temptation to delegate to it in surrender — a decision made far above my head

And I think the examples you provide show very well that trusting it is the act of a fool.

---

<div class="post-metadata">

### Author: ![Velocity](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/velocity/32/18006_2.png) [@Velocity](https://boards.straightdope.com/u/Velocity)
#### Post date: [September 24, 2026, 1:17am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/36 "2026-09-24T01:17:46Z")

</div>

You can’t just trust one AI chatbot for an important question. Ask 3-4 different ones. If they all give the same answer, then it’s probably correct. I routinely run the same query across Grok, ChatGPT, Claude and Gemini.

---

<div class="post-metadata">

### Author: ![dolphinboy](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dolphinboy/32/330_2.png) [@dolphinboy](https://boards.straightdope.com/u/dolphinboy)
#### Post date: [September 24, 2026, 2:08am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/37 "2026-09-24T02:08:34Z")

</div>

Good idea. I’ll start doing that as a double check on important questions.

---

<div class="post-metadata">

### Author: ![DPRK](https://avatars.discourse-cdn.com/v4/letter/d/4491bb/32.png) [@DPRK](https://boards.straightdope.com/u/DPRK)
#### Post date: [September 24, 2026, 2:21am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/38 "2026-09-24T02:21:11Z")

</div>

> [@Frodo](#):
>
> About as far as I can throw the datacenter where it’s running.

Yes, ChatGPT is not “yours”. Even information like the data it was trained on, how it was processed to generate the data it was really trained on, which still tells you nothing about the _model_, is not known by you.

---

<div class="post-metadata">

### Author: ![needscoffee](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/needscoffee/32/1076_2.png) [@needscoffee](https://boards.straightdope.com/u/needscoffee)
#### Post date: [September 24, 2026, 3:30am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/39 "2026-09-24T03:30:44Z")

</div>

Last week the owner of our IT company gave a webinar where he said that right now, it’s considered that AI makes an error about 10% of the time. That’s pretty much what I’ve found, too. Often it’ll contradict itself in the same response. I’ll point it out and it corrects itself. It did that twice yesterday. I noted in another thread where I asked AI for information about a local notorious murder from about 25 or so years ago. It conflated the name of the victim and the killer. It had trouble parsing the article it linked to. And as others have stated, it just pulls info from the web, it doesn’t necessarily judge whether the info is accurate. GIGO.

---

<div class="post-metadata">

### Author: ![Dr\_Paprika](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr_paprika/32/3042_2.png) [@Dr\_Paprika](https://boards.straightdope.com/u/Dr_Paprika)
#### Post date: [September 24, 2026, 3:38am UTC](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319/40 "2026-09-24T03:38:09Z")

</div>

I take anything I am told by a chatbot with some quantity of salt. They do seem to be getting better, but some answers have been really off. They are pretty good for factual questions of mild to moderate importance. For tougher stuff, they are useful in bringing up points and perspectives, but these need further refinement. And there is use in bouncing different answers off other chatbots. But these people who claim to have become business billionaires by trusting the bot in every aspect? Sure! They must be using better bots than I do….

[Previous page](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319.md?page=1)

[Next page](https://boards.straightdope.com/t/how-much-do-you-trust-your-ai-chatbot/1033319.md?page=3)
