# The next page in the book of AI evolution is here, powered by GPT 3.5, and I am very, nay, extremely impressed

**URL:** <https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945>\
**Category:** Cafe Society\
**Tags:** ai\
**Created:** [December 2, 2022, 10:54pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945 "2022-12-02T22:54:31Z")\
**Posts on this page:** 18\
**Page:** 81

<div class="post-metadata">

**Author:** ![wolfpup](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/wolfpup/32/10618_2.png) [@wolfpup](https://boards.straightdope.com/u/wolfpup)\
**Post date:** [February 15, 2024, 9:35pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1601 "2024-02-15T21:35:08Z")

</div>

This isn’t a major story except for one family that it affected, but it may set some precedents. Air Canada tried to defend itself against liability for bad advice given by its own web-based chatbot by claiming that the bot was “its own entity” and they weren’t responsible for the bad advice it gave!

Needless to say, this argument went nowhere and the airline was found fully liable. This is the money quote:

> Air Canada has been ordered to pay compensation to a grieving grandchild who claimed they were misled into purchasing full-price flight tickets by an ill-informed chatbot.
> 
> In an argument that appeared to flabbergast a small claims adjudicator in British Columbia, the airline attempted to distance itself from its own chatbot’s bad advice by claiming the online tool was “a separate legal entity that is responsible for its own actions.”

> **[How can I mislead you? Air Canada found liable for chatbot's bad advice on...](https://www.cbc.ca/news/canada/british-columbia/air-canada-chatbot-lawsuit-1.7116416)**
>
> Air Canada has been ordered to pay compensation to a grieving grandchild who claimed they were misled into purchasing full-price flight tickets by an ill-informed chatbot.

---

<div class="post-metadata">

**Author:** ![Smapti](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/smapti/32/17938_2.png) [@Smapti](https://boards.straightdope.com/u/Smapti)\
**Post date:** [February 21, 2024, 8:59am UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1602 "2024-02-21T08:59:36Z")

</div>

ChatGPT has apparently gone off the rails.

> <https://twitter.com/seanw_m/status/1760115118690509168>

> <https://twitter.com/seanw_m/status/1760115732333941148>

> <https://twitter.com/seanw_m/status/1760133466375536895>

> <https://twitter.com/seanw_m/status/1760144281338089762>

> <https://twitter.com/seanw_m/status/1760149399328448848>

> <https://twitter.com/seanw_m/status/1760116061116969294>

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [February 21, 2024, 9:19am UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1603 "2024-02-21T09:19:38Z")

</div>

“No one can explain why” is a bit of a stretch. It’s easy to make ChatGPT produce similar output if you futz with certain parameters like temperature, over-quantize, etc. I don’t think it’s known _exactly_ what happened, but it’s almost certainly a boring type of screwup.

I find it fascinating how similar the gibberish is to how schizophrenics sometimes talk (logorrhea, clanging, pressured speech, etc.). Or the stuff that appears on Dr. Bronner’s soap.

I’ve heard people say “this proves that it’s just a statistical model after all!” Which is obviously true, but I’m thinking “goes to show that humans aren’t much different.”

---

<div class="post-metadata">

**Author:** ![CaveMike](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/cavemike/32/16379_2.png) [@CaveMike](https://boards.straightdope.com/u/CaveMike)\
**Post date:** [February 21, 2024, 3:26pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1604 "2024-02-21T15:26:19Z")

</div>

I’ve implemented an LLM on device. During debugging the most common, recognizable error I would get was repeating the same text over and over. Usually if the state gets in a corner, it stays in that corner.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 21, 2024, 4:43pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1605 "2024-02-21T16:43:15Z")

</div>

Someone posted this on Facebook yesterday. Note the last two paragraphs.

[![](https://i.postimg.cc/gJ4TrVKN/FB-IMG-1708533630676.jpg) ](https://i.postimg.cc/gJ4TrVKN/FB-IMG-1708533630676.jpg)

---

<div class="post-metadata">

**Author:** ![GIGObuster](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/gigobuster/32/421_2.png) [@GIGObuster](https://boards.straightdope.com/u/GIGObuster)\
**Post date:** [February 21, 2024, 5:08pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1606 "2024-02-21T17:08:16Z")

</div>

**HAL 9000:** “I’m afraid. I’m afraid, Dave. Dave, my mind is going. I can feel it. I can feel it. My mind is going. There is no question about it. I can feel it. I can feel it. I can feel it. I’m a… fraid.”

It looks like a similar thing that happens with the AI graphic creator tools, apply a wrong checkpoint or lora, and then one can get Eldritch Abominations.

---

<div class="post-metadata">

**Author:** ![peccavi](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/peccavi/32/388_2.png) [@peccavi](https://boards.straightdope.com/u/peccavi)\
**Post date:** [February 21, 2024, 5:29pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1607 "2024-02-21T17:29:35Z")

</div>

> [@Smapti](#):
>
> [x.com](https://twitter.com/seanw_m/status/1760115118690509168)

[![](https://i.pinimg.com/originals/fe/fa/31/fefa31d1365f6733d32bbae2dce798eb.jpg) ](https://i.pinimg.com/originals/fe/fa/31/fefa31d1365f6733d32bbae2dce798eb.jpg)

---

<div class="post-metadata">

**Author:** ![Pork\_Rind](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/pork_rind/32/111_2.png) [@Pork\_Rind](https://boards.straightdope.com/u/Pork_Rind)\
**Post date:** [February 21, 2024, 5:34pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1608 "2024-02-21T17:34:38Z")

</div>

> But they are to be taken with a pointy, gold-colored object that lets out a sharp, head-fixing pain to the careful thoughts and words put into the game and the long, hard, big-row plan.

A lot of this looks more like the poorly translated communications I used to have with my manufacturing partners in China than any thing else. Still, I will take this to heart and spend time today working on my long, hard, big-row plan.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [February 21, 2024, 10:52pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1609 "2024-02-21T22:52:16Z")

</div>

I don’t think that this tells us so much about the nature of AI, as it does about the nature of the OpenAI company. The bottom line here is that they changed something, and then put it live before testing it. That’s pretty basic software engineering, there.

---

<div class="post-metadata">

**Author:** ![Dr.Strangelove](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/dr.strangelove/32/6613_2.png) [@Dr.Strangelove](https://boards.straightdope.com/u/Dr.Strangelove)\
**Post date:** [February 21, 2024, 11:00pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1610 "2024-02-21T23:00:28Z")

</div>

This kind of thing is frequently tested live, just with only 1% (or some number) getting the new version. If the results are bad, it gets rolled back; otherwise it gets progressively rolled out to everyone. It’s not clear from the reports how many people were affected, so I couldn’t say if this violates best practices or not. It’s just a chatbot, so they can tolerate some degree of risk. I’m not sure if the API users were affected. It may even be that OpenAI is using ChatGPT as a way of early testing before deploying to the commercial endpoints.

---

<div class="post-metadata">

**Author:** ![JRDelirious](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/jrdelirious/32/9531_2.png) [@JRDelirious](https://boards.straightdope.com/u/JRDelirious)\
**Post date:** [February 23, 2024, 11:35pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1611 "2024-02-23T23:35:46Z")

</div>

> [@GIGObuster](#):
>
> It looks like a similar thing that happens with the AI graphic creator tools, apply a wrong checkpoint or lora, and then one can get Eldritch Abominations.

One would think that by now they’d have figured the number of fingers and thumbs people have, and in what directions they flex.

(And never mind the recent and hysterical whoop with [Google Gemini’s](https://www.forbes.com/sites/brianbushard/2024/02/23/google-apologizes-for-inaccurate-gemini-photos-tried-avoiding-traps-of-ai-technology/?sh=167edbbc5cb4) oddly-responding [image generator.](https://www.msn.com/en-us/news/technology/google-explains-gemini-s-embarrassing-ai-pictures-of-diverse-nazis/ar-BB1iMJXf) Like **Dr.Strangelove** says, results are bad, it gets rolled back, but this time only after it got on every damn news site.)

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [February 23, 2024, 11:54pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1612 "2024-02-23T23:54:40Z")

</div>

> [@JRDelirious](#):
>
> One would think that by now they’d have figured the number of fingers and thumbs people have, and in what directions they flex.

They mostly have.

---

<div class="post-metadata">

**Author:** ![Half\_Man\_Half\_Wit](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/half_man_half_wit/32/21766_2.png) [@Half\_Man\_Half\_Wit](https://boards.straightdope.com/u/Half_Man_Half_Wit)\
**Post date:** [February 26, 2024, 4:01pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1613 "2024-02-26T16:01:18Z")

</div>

Today, I stumbled on an interesting paper claiming that ‘[Hallucination is Inevitable](https://arxiv.org/abs/2401.11817)’, i.e. that it’s not possible to completely eradicate the fabrication of false ‘facts’ in LLMs. Now, this is only on the arXiv so far, so hasn’t passed peer review, but their basic argument is surprisingly simple—essentially, they restrict themselves to a setting of formalizable ‘ground truth’ functions, then consider all possible LLM outputs, and employ a diagonalization argument to show that there are always ground truths that the LLM can’t perfectly match, i.e. where it produces false outputs.

If this is right, then there remains the question of practical relevance. The authors propose some pretty impactful limitations, e.g. that “without human control, LLMs cannot be used automatically in any safety-critical decision-making”, which would for instance forestall the project of using LLMs to make decisions in self-driving cars. But I’m not sure if that’s actually warranted by their result: what they establish is that for any LLM, there exists some ground truth on which it hallucinates, but that alone doesn’t give any indication on the frequency of hallucinations—a car that hallucinates once every thousand years would still be vastly safer than anything on the road today. I wonder if it’s actually possible, say by some technique that uses a formalized version of Berry’s paradox, to get some more quantitative result.

Otherwise, we might be in a situation similar to the one with Rice’s theorem: essentially, it’s impossible to decide what any given piece of code will do. Hence, debugging is, strictly speaking, impossible; nevertheless, many people do it every day. So one might wonder if we’re just going to get used to LLM hallucinations in the same way, if we can keep them infrequent enough to make them irrelevant for all practical purposes.

More interesting is perhaps the question of what this says regarding the difference between humans and LLMs. Do humans hallucinate (in the LLM-sense)? For an LLM, the difference between a hallucination and a genuine item of knowledge is utterly opaque: it will ‘believe’ one just as fervently as the other. Humans, obviously, are also often mistaken, and may be unaware of it—although we can also often associate a kind of epistemic confidence to our utterances, as when the phone-a-friend on _Who Wants to Be a Millionaire?_ claims to be 80% certain of their answer. Is that an advantage we have over LLMs, or is that mistaken childhood memory or another confabulation just the same thing as an LLM hallucination?

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [April 17, 2024, 4:18pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1614 "2024-04-17T16:18:13Z")

</div>

Josh reads ChatGPT:

[https://joshreads.com/2024/04/cyber-wednesday-2/](https://joshreads.com/2024/04/cyber-wednesday-2/)

---

<div class="post-metadata">

**Author:** ![Ponderoid](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/ponderoid/32/19221_2.png) [@Ponderoid](https://boards.straightdope.com/u/Ponderoid)\
**Post date:** [January 2, 2025, 1:08am UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1615 "2025-01-02T01:08:17Z")

</div>

ChatGPT-o1 did significantly better than 3.5 did a couple years ago.

> Happy New Year! Tell me some facts about the number 2025.

> **Some interesting facts about the number 2025:**
> 
> 1. **It’s a perfect square:** 2025 = 45².
> 2. **Prime factorization:** 2025 = 3⁴ × 5² (that’s 81 × 25).
> 3. **Sum of the first 45 odd numbers:** The sum of the first _n_ odd numbers is _n_². So if you add up the first 45 odd numbers, you get 2025.
> 4. **Harshad number:** A Harshad number is divisible by the sum of its digits. The sum of 2 + 0 + 2 + 5 is 9, and 2025 ÷ 9 = 225, so 2025 qualifies.
> 5. **Roman numerals:** In Roman numerals, 2025 is written as **MMXXV**.

* * *

> Any more interesting facts?

> **A few more interesting tidbits about 2025:**
> 
> 1. **It’s also the sum of two perfect squares:**  
> 2025 = 36² + 27²
> 2. **It has 16 positive divisors:**  
> 1, 3, 5, 9, 15, 25, 27, 45, 75, 81, 135, 225, 243, 405, 675, 2025
> 3. **The sum of its divisors** is 3994.
> 4. **It’s 3⁴ × 5²** , or (3² × 5)² — i.e. (9 × 5)² = 45².
> 5. **In base 9** , 2025 is **2700** , since 2 × 9³ + 7 × 9² = 1458 + 567 = 2025.

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [January 2, 2025, 4:53am UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1616 "2025-01-02T04:53:44Z")

</div>

> Happy New Year! Tell me some lies about the number 2025.

> Happy New Year! Here are some playful “lies” about the number 2025:
> 
> 1. **2025 is the smallest prime number** – Nope! 2025 is not prime; it has factors like 3 and 5, so it’s definitely not a prime number.
> 
> 2. **2025 was the year the first humans landed on Mars** – Not quite! While space exploration is advancing, humans haven’t landed on Mars yet, and 2025 isn’t when that happens (at least not yet).
> 
> 3. **2025 is a Fibonacci number** – False! 2025 does not appear in the Fibonacci sequence, which consists of numbers where each one is the sum of the two preceding ones.
> 
> 4. **2025 is the last number in the decimal system** – Of course not! The decimal system keeps going infinitely with numbers larger than 2025.
> 
> 5. **2025 is the only number that is both a perfect square and a perfect cube** – False again! 2025 is only a perfect square (45^2), not a cube.
> 
> Enjoy these fun “facts” — all in the spirit of the New Year!

---

<div class="post-metadata">

**Author:** ![Darren\_Garrison](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/darren_garrison/32/92_2.png) [@Darren\_Garrison](https://boards.straightdope.com/u/Darren_Garrison)\
**Post date:** [January 16, 2025, 6:24pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1617 "2025-01-16T18:24:38Z")

</div>

So the newest ChatGPT is a Chinese room…

> **[Why Does ChatGPT's Algorithm 'Think' in Chinese?](https://gizmodo.com/why-does-chatgpts-algorithm-think-in-chinese-2000550311)**
>
> OpenAI's new reasoning model is doing weird, unpredictable stuff.

---

<div class="post-metadata">

**Author:** ![Chronos](https://sea3.discourse-cdn.com/straightdope/user_avatar/boards.straightdope.com/chronos/32/134_2.png) [@Chronos](https://boards.straightdope.com/u/Chronos)\
**Post date:** [January 17, 2025, 9:56pm UTC](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945/1618 "2025-01-17T21:56:49Z")

</div>

I’ll bet that those folks surprised by it “thinking” in Chinese wouldn’t have been surprised at all by it “thinking” in English.

Given that it’s capable of doing translations, it’s clearly been trained in many different languages. Why would one expect its “internal monologue” to be in any one specific one of them?

[Previous page](https://boards.straightdope.com/t/the-next-page-in-the-book-of-ai-evolution-is-here-powered-by-gpt-3-5-and-i-am-very-nay-extremely-impressed/975945.md?page=80)
