I’ve been using ChatGPT since 2023. I’m using it more now because I’m researching a new screenplay I’m writing. I could do it the old fashioned way and learn it all “fact by fact”, but it’s much easier and more efficient to ask ChatGPT a broadly-worded question and let it tell me everything it knows about a subject. For example, yesterday I asked it to tell me what happens if a dead body washes up on a beach in Big Sur, California and it provided me with three pages worth of detailed information that sounded factual and reasonable.
My question is, how much should I trust ChatGPT? I do spot check a few facts by checking a reputable source, but I don’t fact check everything since that would take as long as researching it all myself. I also know that AI chatbots sometime hallucinate, and I have seen that happen, but it’s pretty obvious when it does, and it doesn’t happen very often these days. So how much do you trust that your Chatbot is telling you the truth without corroborating it yourself?
I recently tried to use Copilot to generate a string of numbers in random order with no repeats and no omissions. Six tries and it was utterly unable to perform this simple task. If a chatbot cannot complete a simple task with clear constraints I have zero faith in its ability to do anything else, like find and provide accurate information.
That’s the problem with chatbots. Despite your assertion that it’s obvious when they hallucinate, in my experience this isn’t true. Sometimes it is obvious, but sometimes they spit out something that sounds perfectly reasonable but is actually complete bullshit.
The more “common knowledge” something is, the better chance you have that the AI is accurate. The more obscure your facts, the more likely the AI is to spit out garbage.
AI is a great tool. As you said, it spit out three pages of detailed information that was exactly what you were looking for. But I personally wouldn’t trust any of it. You need to use the tool properly. You need to verify what it told you. If you don’t, you risk being ridiculed when your screenplay very confidently claims that the letter R only occurs twice in the word “strawberry”. (that particular one has been fixed, but it’s a well known example of why not to trust AI)
I trust it as much as I trust the advice of a dear friend which is smart, but highly yes-manish … and always shares my POV …
so, yeah … not too much on critical “value” things, more on technical stuff (like coming up with a design of a small PV-system, where I can easily doublecheck the numbers)
on more technical stuff, I take anything from it like it was written from a new hire from university … read it benevolent but always with a very critical eye to pick up inconsistencies/errors.
This seems to be the heart of the reason for using it. I don’t think you should trust it at all, but that’s a different question than whether you should use it. Will its imperfect information get you where you want to go? Yes. Will you get some wrong information? Probably, yes. Will it matter? Well, it would matter to me, but most people watching the screenplay once produced will, at best, ask ChatGPT if they think you’re inaccurate, by which point your screenplay will be part of its database.
I think of it like that one guy you know who is useful at times but most of the time just spouting bullshit. Like you know some of it’s probably true and it’s entertaining to hear about, but you’re never going to fully believe anything that guy says.
I trust it less and less each time I use it. The same overconfidence tells that identify human bullshitters have become even more pronounced than they were when I first encountered AI chatbots.
ChatGPT has one of the lower honesty rates. OpenAI seems to focus more on having their AIs able to solve hard problems (math, logic, etc.) than ensuring that the agents understand source evaluation and chains of evidence:
That said, if you go up to the top models (Sol vs. Luna), it’s going to be better.
I use Gemini at home and (mostly) Claude at work, so I don’t have any experience with the ChatGPT UI to know how much detail it provides about what the agent did. If you can see its web searches or at least see that it did perform a web search, then that should give you some confidence.
But, really, what you want to do is to create a custom GPT and give it instructions about how to perform research, so it doesn’t have a choice about how to retrieve information.
One issue in modern times is that Google stopped preferring educated and citation based research for “any other person’s popular blog”, so even when a model does perform a web search, the topic might be one where the woo page rank is higher than the scholarly. (Technically, I think that ChatGPT searches using some other search engine than Google but smaller options are going to have just as hard of a time avoiding woo content. The issue is probably largely just that the mass has become too great to avoid.)
So, really, unless you provide it with a means of identifying between scholarly material and woo material, you’re more likely to get thrown off by that than by the model itself hallucinating - so long as it is actually performing Internet research and not relying on its internal memory.
I use ChatGPT, and I’m pretty happy with it. If it gives an answer that may be in doubt, it will say as much. I think that’s an improvement I like. I view it as a reliable source but not infallible, and that is a good approach for me to have.
Its been very helpful, but it makes mistakes. I’ve asked it the same question at different times and gotten different answers. Its made basic mistakes.
AI is still in the early stages. It isn’t like it magically appeared with GPT-4 and now its at its peak 3 years later. Its a slow process of making it more competent thats been going on since the 1950s and theres a lot more improvements to make in the decades ahead.
its very helpful for me, but it gets things wrong often enough that its a serious concern.
Having said that, Gemini notebook has been extremely good for me since that is searching through technical manuals. As was mentioned earlier, a lot of cites on gemini are just reddit posts.
Thanks for all the replies. I’m not using a chatbot to write my script, although some people have done that, I’m using it to provide information I could dig up myself. If it’s something critical to the story, I will research it, but for something that’s procedural and well documented, I will let if tell me what it knows and decide if I trust it or not. I never said it was perfect, and that I blindly accept everything it tells me. However, I’m suggesting it can be a productivity tool, at least in some circumstances, and that’s good enough for me.
“Whether to trust a chatbot” isn’t (or at least shouldn’t be) a binary yes/no thing… it’s like asking “How much do you trust this brochure?” So many factors are involved:
Your question (is it some basic thing, like “Where is the Eiffel Tower?”, or some more obscure factoid that is not well-represented in its training set)
The way you prompted that question
The model in use (like Gemini Flash 3.6, or Claude Fable 5.1, or OpenAI Astra) and the effort level it’s set to (like “thinking off” or “maximum”)
The “harness” in use, whether it’s your chatbot’s “Research” mode or a purpose-built one for research, like Gemini Notebook (as Wesley_Clark said). Both will be better than a basic prompt in a chatbot, and Notebook in particular is really good at verifying its own research and providing in-line citations, and it generally prefers more trustworthy academic sources. It is still not perfect, so always verify.
Whether your question would likely get entangled in someone else’s profit motive (i.e. whether it is likely to be a target of GEO, the LLM version of SEO spam)
Whether the question you’re asking even has “clean”, known-good sources that pre-date AI slop. Many sources, even academic ones now, have been “polluted” by AI slop and hallucinations, and it’s becoming more and more of a vicious cycle of AIs quoting other AIs to make more AI slop, and it’s quickly exceeding our collective capacity to actually review any of it for truthiness. We’ve run into the literary equivalence of running out of “low-background steel” to act as a control group, and so facts and texts in general are becoming less and less verifiable by anyone, machine or human. Old books are more valuable now, and AI training itself might prefer texts published before 2021 or so.
All that said… I trust LLM research about as much as I’d trust your average college undergrad literary research project, which is to say: Not that much, but it’s often a good starting point. More than necessarily telling me all the answers, it provides a good framework for what to think about and what the relevant questions to ask are, and from there I can do more research (LLM-assisted or otherwise).
I’ll usually start with the research mode, then ask it to spawn some sub-agents to double-check its own work against the original sources in a fresh context (sometimes you’ll have to get the full text PDFs on your own first, e.g. via a library subscription), and then do some more manual checking using kagi.com (since Google is all AI spam now). I’m still never truly sure of the truthiness anymore, unless it’s something I am able to independently verify on my own.
TLDR LLMs are not innately trustworthy, but they are useful, and can be made more trustworthy through the application of certain techniques and tools — though never perfect. Further, verification itself is an increasingly difficult task even for humans, given the amount of AI-generated spam in source materials.
All these similar AI threads make me wonder if we could benefit from some collectively-edited “How to AI” guide, with clear instructions and examples for common Doper use cases (research, computer help, editing images, producing reports, coding, etc.).
The combination of model X effort X harness X prompt can make a huge difference in output quality, but they are generally not very well explained by the AI vendors themselves…