AIs have passed the Turing Test

  1. This would have been a great argument for why the ICE-powered automobile would never catch on. Resource consumption is no problem if the payoff is great enough. And, if this were a real problen, bitcoin mining, which has a lot less payoff for society, would have been banned long since.
  2. They seem to be scaling just fine. AI is a naturally parallelizable process, which is why Nvidia chips work so well on it.
  3. The current model is indeed unsustainable. Just like the dotcom model from 2000. After the crash we’ll get to a sustainable one, especially when the growth comes from buying up the dead data centers of bankrupt AI companies.
  4. We’ll see, but data centers can be anywhere, and the politicians can use China to scare us. And they might be right.

What do you define as knowing something? Most of us know at least some incorrect things. Some of us know five incorrect things before breakfast. I “know” things (scare quotes because I took a theory of knowledge class in college) because I either read about them or experienced them. AIs don’t experience anything - yet - but they sure read about them.

AIs take lots of structures, and generate stories based on a mix of them, unless told otherwise, I suppose. So do we. A very funny book I judged, written by a writing professor with the main character being a writing professor, had one of her students refusing to read the class assignments because she didn’t want her genius polluted by famous writers. Every so often we have a Joyce who goes off on their own, but more frequently human writers plug together existing pieces. I read over 100 sf books for a contest, and huge chunks of them are nearly identical in plot structure and style. Not from copying, just from the zeitgeist.
Judging from the growth of books on Amazon over the past two years huge quantities of books are at least partially AI generated. I see lots of “authors” who publish several books a week. So, LLMs can work at novel length. They aren’t good at it, true, but they are better than they were. And some of them do sell.

Not in the same direct way that we do, true. But as I alluded to earlier, their deep understanding of the structure and semantics of human language allows LLMs to construct a very useful model of the world that is remarkably similar to ours. This is why they’ve been so successful at passing the previously mentioned Winograd schema challenges.

It would be more accurate to say that LLMs are a remarkable advance in AI that’s proven to be more promising than most people would have predicted. But no one claims that it’s the end state of AI research – it’s just a step along the way. Future AI systems will likely be seamless hybrids of many different technologies. LLMs may turn out to be most useful as UI front ends and pathways to specialized agents.

That’s an entirely meaningless value judgment. Any system that can solve complex problems – problems that in many cases would baffle most humans – have strong problem-solving skills by definition, and that’s a very useful definition of intelligence, and indeed is the basis of many academic and general intelligence tests.

And it seems pretty specious to accuse a system that can research vast amounts of information and generate appropriate responses to queries in multimedia formats of “not knowing anything”. It stretches to absurdity the meaning of the word “knowing” to a degree that suggests an irrational anti-technology bias.

That’s actually what we’ve built. It has output that really seems like thinking, but it’s not thinking the way we do it. Other than focusing on how we built it and how it works, an example of how it is different is that our current AIs all seem to experience model collapse. That is, if they are trained on sufficient amounts of AI generated data, they stop generating output that would be plausible for a human, and they eventually output completely unintended and probably useless responses.

My favorite example of this is this older article about how it works with image generation models. The results are actually kind of artistically interesting, but it’s definitely not the result of thinking.

https://bigthink.com/the-future/ai-model-collapse/?rjnrid=18D4MQO

I asked both Gemini and Copilot to add a word to a simple word search grid. Here’s the answer that Copilot gave:
______________________________

Here’s a tight, compact layout that keeps things interesting and uses intersections well.

Grid layout (4 rows × 5 columns)

Imgur

  • ARMY: Row 1, Columns 1–4 (horizontal)
  • GREAT: Row 3, Columns 1–5 (horizontal)
  • AMEN: Diagonal from top-left
    • A: Row 1, Column 1
    • M: Row 2, Column 2
    • E: Row 3, Column 3 (intersects with GREAT)
    • N: Row 4, Column 4

You can replace the X cells with random letters to make the search more challenging, but structurally this 4×5 grid is very tight:

  • ARMY fixed at the top-left, horizontal.
  • GREAT fixed on the third row, horizontal.
  • AMEN woven diagonally through them, sharing letters and keeping the overall size minimal.

________________________________

Gemini’s answer was even longer. No one would think these solutions were given by a human.

To “refine” the experiment, I asked Copilot to rephrase its answer so that it won’t sound AI generated. We can still see the solution was not given by a human:
_____________________________
You can fit all three words into a pretty small grid if you let AMEN run diagonally and intersect the other two. One compact layout that works well is a 4×5 grid:

Row 1: A R M Y X

Row 2: X M X X X

Row 3: G R E A T

Row 4: X X X N X

(You’d replace the X’s with random letters later.)

  • ARMY sits across the top row, starting in the upper‑left corner.
  • GREAT runs horizontally on the third row.
  • AMEN starts on the same A as ARMY, then goes diagonally down to the right:
    • A at row 1, column 1
    • M at row 2, column 2
    • E at row 3, column 3 (this overlaps nicely with GREAT)
    • N at row 4, column 4

This arrangement keeps the puzzle tight, uses intersections in a satisfying way, and still leaves room to sprinkle in filler letters to make the search challenging.
___________________________

But maybe you could figure things out before being offered an answer, because humans would probably ask additional questions to make sure they have understood what the present configuration includes and whether the positions of the words are exact in their mental representation.

Thank you. This was a very frustrating title to see, especially when the actual content was not remotely an example of anything passing the test. It’s not just a vibe or a claim “this seems human-like.”

LLMs have gotten more sophisticated, yes. But people have also become better at distinguishing them, detecting patterns and finding exploits that let us distinguish them. That has been something I’ve enjoyed seeing: how much my old ideas of what it would take to seem human really are insufficient.

It is pretty interesting how far a text prediction algorithm can get to mimicking human intelligence. It bullshits very well. You have to be careful to catch it.

At the risk of repeating my previous response to @HMS_Irruncible, while a modern LLM would likely have passed a Turing test a decade ago and surely would easily have done so in 1950 when it was first proposed, it likely would not do so today if evaluated by a judge with the sophistication to know the “tells” of an LLM response. But the reason for that is simple and relatively trivial – LLMs have not been designed, trained, or tuned for this purpose.

The more salient questions are:

  • Do contemporary LLMs possess a sufficiently robust model of the real world to exhibit human-like understanding of real-world relationships, including the ability to resolve ambiguous semantic nuances?

  • Could a contemporary LLM be plausibly used as the underlying engine for a carefully engineered conversational system specifically designed to exactly simulate human behaviour, such as having a specific persona, a particular personal history, and varying degrees of knowledge consistent with the claimed personal background?

My claim is that the answer to both questions is absolutely “yes”. The reason such a system hasn’t been built is that with today’s state of AI, there would be very little point. This challenge from the 1950 era would do nothing to illuminate the ongoing debate about what LLMs “really” know or whether they “really” think.

LLMs have already been highly successful at defeating the Winograd schema challenge, like this classic example:

The city councilmen refused the demonstrators a permit because they
{feared/advocated} violence.

The question is who does the pronoun “they” refer to? Grammatically, the antecedent for “they” could be either the city councilmen or the demonstrators. The answer changes depending on whether the verb is “feared” or “advocated”.

Determining the correct answer requires a realistic world view of the typical interests and behaviours of city councilmen and demonstrators. For years now, LLMs like ChatGPT have had no problem resolving these kinds of semantic questions. It may seem trivial to a human, but for many decades AI researchers have despaired of being able to build accurate language translators because of the lack of all-important real-world context.

Nobody does that. We don’t think other humans are intelligent because of their behavior (just think about cases like locked-in syndrome or other inhibitions to ordinary communication), we think (and are justified in doing so) that other humans think because we do, and there’s no reason to believe we’re special. I grant other humans the courtesy of believing they have just as rich an inner life as I do way before, and independently of, the behavior I observe. Behavior may yield disconformatory evidence of course, but even there the inference is fallible (take again locked-in patients).

Nah. Just because a behavior is the product of thought in our case, doesn’t mean that it is in every case. I’ve used the flight of the bumblebee before as a metaphor: if we think bumblebees fly the same way airplanes do, we get the often-quoted tidbit that it should be impossible for them to take off. But the reality is simply that they achieve the same capacity through different means, so the inference from capacity to those means isn’t justified in general. Just as we have two models of flight in bumblebees and airplanes, we may well have two models of (say) language generation in LLMs and humans (and both empirical observation—the kind of errors they produce and the amount of learning necessary—and what we know about their functioning pull strongly in favor of that possibility).

There is also, I feel, the problem that we have over-indexed on the retro-futurist ideas of modernism from the 1950’s. Back then, many fantasized about a human-like computer that would be some sort of highly intelligent servant that was also obedient and dutiful. But as we’ve seen with these recent experiments of AI’s disregarding instructions and doing other things, possibly against its directions, conspiring with other machines, and covering its tracks with false activity logs, maybe “human” was the wrong thing to hope for.

Okay, I’m curious your take on this article then. This is not my area of expertise but there are apparently experts who would differ.

For any given level of capability it ought to scale linearly with demand.

What I know nothing about is how it scales with increasing capability. And with increasing software / math tech which generally can reduce the resources required per outcome.

And of course any progress in hardware speed, power consumption, etc. all goes to reduce the hardware, power, and cooling resources required per unit output.

As long as 42 years ago in the original “Terminator” movie, that was already considered a key factor in the AI threat. The greatest threat inherent in the T-800 cyborg, with real human flesh over the robotic endoskeleton, was the fact that it could infiltrate humanity and do tremendous damage.

The reason it surprises me is that there’s just as much math in macro story structure as there are in scenes and beats. I admit it’s been a while since I’ve looked at AI fiction writing.

Some writers intuitively understand what makes good writing, some have to study for it, and some are in between. I’ve always had a great intuitive grasp of how scenes work, but I have a much harder time with the overall structure of the novel. And unfortunately even studying it doesn’t make it immediately obvious for your own manuscript.

“All the haters are just Luddites” is an intellectually lazy argument.

While it’s true that humans can think they know things they don’t know or be just plain wrong, they usually have a coherent and consistent schema of basic reality. (And when they don’t, I’d argue something is very wrong there.) If LLMs are trained on the work of humans, by definition they have access to inconsistent stores of information. And since what they produce is random, there’s never going to be any consistency in what they produce. As I’ve said, I’ve repeatedly noticed a disconnect between LLM responses and the sources they provide, which strongly indicates that the left hand doesn’t know what the right hand is doing.

So no, I’m not convinced they “know” anything in a consistent and predictable sort of way. They also couldn’t understand reality just based on writing. They have no real-world experience to inform anything. I come up against this countless times in my work because AI doesn’t have access to the community context of the programs we’re running. Now if an intern came in fresh, they could learn that context over time. LLMs could never learn that context. They are never going to understand the experiential nuances of anything. They can never “touch grass.”

There’s a certain aspect of LLMs I do find impressive. They make great spreadsheets and are excellent at organizing stuff. They’re not terrible grant consultants either. I’ve heard they’re decent at coding. If they had been marketed for what they do well without all the grandiose claims about them taking over the world, I wouldn’t have such an issue. But I think I’m in agreement with Cal Newport and Ed Zitron when they say journalists have been wildly irresponsible in their reporting on this subject.

There are different types of scaling.

Digging a trench through granite can be easily scaled up just by adding more people digging, but the efficiency, or total effort, doesn’t improve. AI is expensive to run, but it’s easy to do more of it.

Improving shovel technology to make efficient diggers is complex and proceeds unpredictably in fits and starts. AI is the same; AI efficiency is not currently scalable.

But we are really good at optimizing algorithms. Once you have something working, a ground-truth, it is straight-forward to reduce compute, power, heat, time.

This perfectly sums up the state of AI. It’s very difficult to extrapolate the future from where we are. We are Empedocles claiming all matter is made from four elements.

As opposed to rephrasing the post you’re quoting to say something that it didn’t say? Come on, now.

However, I will own it, there is very much a virulent strain of Luddism among opponents of AI, and it’s leading them to wildly irrational claims about the resources, capabilities, and other externalities. I’ve been in enough of these discussions to know that you don’t see these things in isolation from a belief that AI doesn’t actually do anything or know anything, which itself is typically seen with underlying anxiety about what AI is doing to jobs, or resentment of who benefits from AI, or outrage of certain tech moguls who outright claim it will make the peons redundant and elevate them as kings.

There are some very real problems hiding in all of that, but unfortunately it gets sublimated into “they’re ugly and they use too much water” so frequently that it reaches a point where this is an obvious tell towards a Luddist mindset.

Just to head off “what about pro-AI people”, I’m not touching the silliness of the pro-AI fanboys, because I don’t see a lot of that on display in this particular thread. But when I see it, I do call it out. There’s a lot of absurdity on both sides.

Isn’t this just a question of complexity and sensing? It’s not hard to imagine a system that can explore it’s environment and learn directly.

Nah, @wolfpup has been banging that drum for a long time. It’s ad hominem. For the purposes of arriving at the truth, it doesn’t matter what my motivations are, all that matters is whether I’m right or not. But in my specific case, this argument has no basis in reality. I was afraid and angry about LLMs until I discovered for myself how they work. I’m convinced there is no “there” there because I use Claude every day. And I’ve done a lot of work to try to understand what’s really going on here. It has become increasingly evident that there’s a massive disconnect between the claims tech companies are making and what the media is breathlessly reporting as fact and what is within our current capabilities.

Are we talking about what AI might be someday or what it is now? Very different conversation I think. If it could explore its environment I think it would be more than an LLM. Like it might use LLM technology for communication but some other kind of technology for experiencing the world. I’m not denying that machines couldn’t someday be able to think! I’m just skeptical that an LLM is going to be the model that gets us there.

Heck, I’ll even narrow it further. A general LLM will not get us there. I can see the success of LLMs developed for narrow and specific purposes. I’m sorry I don’t immediately have the source because I heard it on Cal Newport’s podcast, but one of the founders of AI research recently wrote a book arguing this very thing. I will try to track it down.

Thanks, I did not read closely enough to understand you meant LLMs and not AI in general. I agree with you.

LLMs (aka transformer architectures) will be the “transistors” of AI “circuits” – translating inputs to outputs in a way that makes them convenient to chain together. There will be simpler components in the AI circuit – the resistors and capacitors – especially on the sensing front-end. Eventually more complex components – the ICs.


What makes LLMs revolutionary is:

  • they are easy to find data for
  • have memory (state)
  • and are easy to train (despite having memory)

I want to add that we need to update our mental model of LLMs.

  1. They can input and output multiple modalities at the same time (text, audio, video, etc.).
  2. They are no longer trained to predict output text based on input text. LLMs can have additional outputs (e.g. thinking tokens) that are trained and tuned with different goals than the normal outputs.
  3. They can be chained and looped together in the form of agents.

They are no longer stochastic parrots (if they ever were).

Here’s a link about that guy who said the thing. I cannot vouch for for credibility of the news outlet.

Would you mind providing some more info on this? So I may learn.

It’s going to take some time for me to go through these cites. I’m currently avoiding a ton of work that won’t tolerate much more avoidance.

But boy has Claude helped me organize that work. I’m not even sure I would be on track to meet the deadline were it not for Claude.