2026: Claude vs ChatGTP vs Whatever

I’ve read Vonnegut…

Noticed that.

Nah, Claude’s icon looks either like a brilliant supernova or the result of dropping a raw egg on the kitchen floor (and trust me, I know what that looks like). GPT’s icon looks like a confused Mobius strip that got entangled in itself. Gemini ripped off its icon from the American Iron and Steel Institute (the logo originated with US Steel). As for Grok, they should just be honest and use a swastika as their logo. :face_with_diagonal_mouth:

I’ve learned today of an interesting and simple illustration of how LLMs work. Ask one to give you a number between 1 and 10, and it’ll probably give you a 7. Ask again and you’ll likely get a 3.

I have tried this with Gemini just now and it was true for me.

It’s not a glitch, the model is simply drawing on the most common human answers to that question. 7 is the most common answer, 3 is the second most common.

This demonstrates that the models give you an average answer to so many questions. This shows that LLMs don’t create distinctive ideas. I think a lot of companies that turn to AI will start to behave like each other. At the very least users need to prompt very carefully.

PS Anthropic just released Claude Science today: Claude Science beta | Claude by Anthropic

It’s a specialized app for (primarily) data science in biomed & pharma right now, with Python, Jupyter, and visualizations and citations etc built in, along with a grasp of LateX. It comes with connectors to arxiv and a few other resources, but not yet the scope of Google Scholar.

LINUS: What’s so bad about being average?

LUCY: Because you’re capable of doing much better!

LINUS: That’s the average answer!

Hehehe, I have been running gemma2:2b locally for a bit now, and it outputs 7 to the first question in a session to pick an integer very consistently. On the first repeat of the question it is likely to pick 3, but sometimes it picks 5. Even though I understand how it works, I was surprised to find out it works well on the open weights model, as well.

But phi:latest? No, it refuses to pick a random number for me. It decides to instead give me a dissertation as to why it doesn’t pick random numbers. It also seems to give only sane responses to “What should I do if my toppings keep sliding off of my pizza?” No mention of putting parchment between the crust and the toppings, and no suggestions such as placing little books on top of them while cooking them (yes, gemma2:2b did suggest both of those).

You’d think phi:latest would be a better model, generally. But kind of like the gemma2:2b model, if you ask it to do things it can’t do, it may just give you a treatise on what it can do, instead. Oh well, it’s just costing me cycles on my 3d card.

I used ChatGPT for a few years but have switched more to Claude over the last year and have really enjoyed it. My sense is that Opus has an analytical and writing quality somewhat better than its competitors.

All this is doubly true for Fable 5 which is just in a class of its own. As you may have heard it was taken down because of the Trump administration but is back and will be available in the $20 plan for several days. I have been taking it through its paces discussing public policy issues which I know something about and , man, the quality of its analysis is sometimes stunning, you really get the sense of interacting with something intelligent. I just hope that as inference costs come down, they can figure out a way to make it available on the basic plan again.

Fable 5 is also supposed to be fantastic for coding and before it goes away I plan to use it to design one or two simple games and maybe an interactive simulation of some kind.

I have a slight preference for Claude’s style over ChatGPT (free versions in both cases) but IIRC even the free version of ChatGPT gives you its most powerful model, though for only a limited number of compute cycles before it downgrades.

One thing I noticed recently about ChatGPT that I quite like is its ability to correlate the things we’re discussing with previous unrelated conversations. It may say things like “based on your interest in …” or “your concerns about …” it tailors its responses to opinions or interests I’ve expressed elsewhere. I think it’s a very cool feature that makes it seem more like a friend than a robot.

I take it that you don’t have concerns as expressed over here:

Frankly, no.

I’m an advocate of personal privacy for many principled reasons, but I’m also not paranoid about it. If an LLM says something like “here’s why I think this new information is relevant to what you mentioned before …” I think this is genuinely useful. I was a little startled when ChatGPT knew roughly where I lived based on my IP address, but that still didn’t bother me. I don’t see how this can be harmful, except maybe in terms of targeted advertising. So what? I am not a fugitive hiding out from les federales. I’m just a humble pup minding his own business.

I’m with you but … :slightly_smiling_face: … I did ask Claude … and the point it makes that does give me more pause was this one:

Asymmetry of who’s looking and why. “Easily knowable” assumes a neutral or friendly looker. The people who actually go looking for aggregated profiles are disproportionately doing it for leverage — a competitor, an opposing counsel, an HR issue, a scammer building a pretext. The information doesn’t have to be secret to be dangerous in the wrong hands at the wrong moment.

Yes there are stated controls and I am even naive enough to believe they really are there … but breeches will inevitably occur. An AI powered scammer empowered by a decent enough psychological and fact laced profile of who I am, who my relationships are, and my financial information? Using it to not only scam me but perhaps to empower scamming others? That’s scarier to me than ads.

I’m not yet changing all my settings to max out privacy, but I do now better appreciate why I probably should. Really.

I asked Chat GPT and my first response was 4. Second was 9.

And that’s before some future despotic government decides people like you are the ones they’d like to persecute starting today. For whatever definition of “like you” they’d care to name.

As we have seen both in other countries and now in the USA, despots choose their targets not on the basis of actual harm to society, but for their convenience as a PR exercise to fire up their base.

For LLM use, I use Claude. Started using it for political reasons but really I just like it better than Gemini or ChatGPT. My only complaint is that it doesn’t do images so if I need something rendered as part of an LLM response, I’ll use Gemini. I don’t use it super often and my questions are often light coding-style stuff (really more like “check this config file and make any suggestions”) or sometimes a work related question (“what would you budget for this task”) and of course I check the answer against my own experience but it’s usually a reasonable response. I don’t really use LLMs for chitchat, life advice, etc.

For images, I mainly use local AI but online I use Google’s Flow suite (Nano 2 and Nano Pro). I find it much better than working through Gemini for doing “art”. I also have a sub to Midjourney and really like their outputs but I haven’t used it in the last month or so (just haven’t needed to) so should cancel for now and save the $10/mth.

I think there are two different and distinct issues here. The ability of an LLM to search and process vast quantities of information at seemingly miraculous speeds is certainly a capability that’s subject to abuse by malicious actors. That’s a given and that is unfortunately the world that we live in, and there’s not much we can do about it.

The point I was making is that the apparently new capability of ChatGPT to maintain a profile of my interests aggregated across different conversations is impressive and seems to me to be a net positive that allows it to tailor its responses to areas that are most relevant to me. In that way, as I said, it seems to be acting more as a friend who knows me rather than a machine providing robotic responses like, say, Google. Claude can probably now do this, too, but I don’t have as much of a history there yet.

I definitely get that positive. It is why I am not rushing to block as much as possible: I experience that as utility.

I am however definitely acknowledging that I get that utility at some margin of risk above what is otherwise easily knowable about me, more than just targeted marketing.

Is there an alternative to NotebookLM? I love that program. I upload a variety of PDFs and then search through them. its nice if I have a book or manual I want info from. (nevermind, I checked and its called claude projects).

I personally stick with Gemini and GPT since those are the ones I have a $20 subscription with. But of them I notice Gemini makes more basic math mistakes. I don’t know if that has changed or gotten better, since I don’t use AI for math very much.

For NotebookLM, the “Research” modes of Claude and ChatGPT should be similar enough for basic uses, maybe? It’s not exactly the same but you can provide PDFs and web sources and tell them to research more and summarize and chat with them etc. But also, why not just use Notebook itself? They renamed it recently but it still works.

For math, if you mean basic arithmetic, try using an agent harness (Claude Code or Google Antigravity or OpenAI Codex CLI) where they have more freedom to use local tools, usually Python scripts, to do the calculations. The regular chatbots can do that too, but you get more explicit control with the harness. Either way, if you just explicitly ask it to write a Python script to do the math, it should be more deterministic. LLMs aren’t great at arithmetic but they can easily write code that IS, in a script or spreadsheet, that they they evaluate the results of. Claude is very good about knowing when to do this automatically; with the others you might have to be more explicit.

(Anecdotally, over the last few weeks I’ve been using Claude Code to do arithmetic heavy hobbies… it’s making me a 3d video game, making new models for my 3d printer, redoing PDF layouts, etc. The model itself can’t necessarily do all that in its head, but with a harness and tool calling it becomes trivial for it.)

If you mean higher level math, like proofs and such, I don’t have enough math background to say anything useful, but here’s a transcript and discussion about how Terrence Tao was working through a problem with ChatGPT… https://news.ycombinator.com/item?id=49010345 (I don’t know either the person or the problem)

Claude for work. Chat for personal life.

Claude cowork is next level. Literally doing work for me on my computer while I was eating dinner checking in on the progress was not something I knew existed this world.

As I mentioned in another AI thread, I was selected to participate in a trial from Yale University of an AI called “ourai.” I don’t know why I was chosen, it seemed random, I did answer some qualifying questions (bland and pretty non-specific) before starting. The goal of the study seems to have to do with how much a person’s political policy views might change based on interactions with an AI. While the study is on-going, I have free use of the AI for any queries of my own. The first few weeks the testers were directed to make queries on a particular subject (and they gave sample questions we could use). This last week we were directed to make 10 separate and unrelated queries of our own devising.

I have used it only for queries, that seems to be the only use for which they have opened it up. I had not used AI before (except in Google searches as it comes up automatically) so I don’t know how it compares to others. I do have reservations about it, specifically:

Responses come with references, numbered links within the responses so as not to break of the flow. The only way to see what the references are is to click on the numbered link.

Most of the sources seem like reasonable places to look for answers, once in a while a source seems dubious to me, mainly in that it seems like it might be pushing a particular point of view. I haven’t done any systematic survey in the responses I have gotten to see how many are like this, or how bad they really are.

It seems there are never more than 6 sources referenced in any one answer. Sometimes the numbers go up to 6, but there might not actually be that many sources linked to, like recently source #5 was missing in a response. I asked about that but got no answer (it was one in a multi-part question asking about its previous response). I suspect it is searching until it finds 6 sources, and then stops looking. I don’t have a sense that it is evaluating all the potential sources and picking the best ones. Some sources it uses are certainly better than others.

On one query I made it completely misunderstood the question, even though the query was very clear, and was re-stated in brief in the last sentence of a moderate-sized paragraph. It apologized when I pointed this out, and gave me a response to the actual question, but I found this unnerving. I don’t know if this sort of thing happens in other AI models.

I haven’t detected political or policy biases in it, during this very limited trial. It seems able to discuss nuanced topics reasonably well. But if I were to rely on it for any real-world application, I would certainly check all the sources that it uses, and maybe ask it to answer the same questions using different sources from the ones it used the first time. Now that I think of it, I believe I will try this on one of the trial-directed queries.