Should AI splooge be treated the same as copyrighted materials on the Dope?

The notion that science fiction writers originate ideas that then filter into the mainstream is an example of slop that proliferated well before AI. They may very well popularize ideas, although virtually always in ways that are unrelated to reality, but those ideas virtually always come from first seeing them in the real world. Jules Verne said “I always took numerous notes out of every book, newspaper, magazine, or scientific report that I came across.” Nothing has changed since.

This is absolutely not true. Large Language Models are called generative AI precisely because they generate brand new content based on what they’ve learned about how the world works and how to describe it in language, and to distinguish them from other types of AI such as predictive AI. At its core, the learned patterns and concepts and the ability to express them in language is similar to the way humans learn and create content.

Exactly.

That said, how AI output should or should not be used on the SDMB is a matter of board policy, as @engineer_comp_geek described earlier.

LLMs do not store any copyrighted work of text wholesale. It’s not how they work. They’re trained on huge amounts of data, much (probably most) of which is copyrighted, and it forms a set of statistical weights that makes associative links between that data. Certain very famous phrases like “to be or not to be” might make it into the data wholesale but that’s just because those words are so strongly statistically linked by repetition in human writing that they’re highly likely to be reconstructed rather than the LLM having a full copy of Romeo and Juliet stored somewhere. The original training data is lost after the training, only the statistical weights remain afterwards. I would argue that no copyrighted works exist in LLMs, only a synthesized set of statistics about copyrighted works.

And I think this a pretty critical distinction because humans work in a similar way. Your brain and your understanding and the words you choose to speak is influenced, too, by all of the works you’ve seen and read in your life. Not exactly mechanistically in the same way, but I think it’s philosophically close enough to say that the idea that having been trained on copyrighted works and then synthesizing your own answer that does not include the original material but that includes your knowledge of them is something both humans and LLMs do.

I do not think it’s reasonable to say that LLM output is a copyright violation. And if you look at the language of some of the people proposing this idea, you can see they’re extremely biased against LLMs, so this seems like a roundabout way of trying to ban any quotes or content from LLMs from the board rather than a serious discussion of copyright.

That said, you can have whatever policy you want on how to present LLMs to the board. I only bring LLM output to threads about AI, and if I’m quoting a full response I label it with a details box so people can choose to read it or not. If it’s a short snippet from a response rather than a verbatim response, maybe a quote box would be appropriate. Or if it’s images or video it’s clearly labelled. I think disclosing that something you’re including was written (or created, in the case of images, music, video, etc) by AI in the same way you might do if you were quoting a friend you had a conversation with or you were quoting someone else’s tweet you found funny or perhaps an encyclopedic source of information which is also a synthesis.

Factual answers might be a slightly different category because it can be more like quoting an encyclopedic source, but since the prompt heavily shapes the outputs that’s not quite right either. I would say in that case, LLM generated responses to factual questions should be discouraged in some place like FQ and clearly labelled. Discouraged simply because the OP would have the option of going to an LLM themselves, didn’t, and copying an answer in that way is basically doing the equivalent of “let me google that for you” which is also discouraged in FQ.

I’d say that an AI quoting “To be or not to be” is definitely not due to it having a full copy of Romeo and Juliet.

That said, there are enough references to enough parts of Shakespeare’s plays in enough sources, that I wouldn’t actually be surprised if an AI did have some of the more popular plays memorized in their entirety (just like some humans have memorized the plays). But of course, that’s still not a copyright issue, because Shakespeare’s plays are all in the public domain. And for any work that’s not in the public domain, it’s unlikely that it’d be quoted enough to get memorized like that.

Haha, excuse me, Hamlet. I got my wires crossed because that wasn’t the part of my post I was trying to be careful about.

It’s actually easier to see in image generation in terms of what an AI knows versus what it’s synthesizing. If you ask it to draw the Mona Lisa, it’ll probably get close - recognizably the Mona Lisa but some errors / differences here and there. That’s because it’s one of the most famous visuals in the world and appears all over the training data. So the statistical weights are so strong that it almost gets there.

But if you ask it to draw less well known / well represented things, you get more and more differences introduced that make it further from the real thing. You ask for a moderately famous painting and it probably draws something roughly in the right ballpark - similar style, probably the correct subject, whether it be an old man or a young woman or a fruit basket, but noticibly different from the original.

Which shows that it doesn’t store pictures, exactly, it stores statistical weights. Which is the same as it does with text, but it’s easier to immediately compare in picture form. By the time you give it a novel picture idea, like a novel argument or just a particular set of words in a prompt, what it will come up with will be original synthesis that likely is not that close to anything that appears in any existing textual or visual work that was created by humans.

I would be surprised. @SenorBeef is essentially correct where he says what I was trying to say earlier, but he says it more eloquently and completely:

There are some exceptions, like memorization of passages that are found very often in the literature, and “mosaic memory” of fragments of academic papers (but not the papers in their entirety) which, again, is a close approximation to how human learning and cognition works.

What I find especially profound is the fact that, by acquiring high proficiency in natural language, LLMs capture the human view of the world that is intrinsically encoded in the grammatical structure, syntax, and semantics of human language. I think it’s fair to say that when an LLM responds like a human, it’s at least partly because it’s learned to think like a human, and certainly not because it’s regurgitating something in its storage.

I think the bottom line here with respect to the OP question is that:

  • Responses generated by an LLM should be regarded as original material

  • By law, that material is not subject to copyright

  • But it’s subject to SDMB posting rules, which at the very least requires attribution, and is also subject to a preference for not posting lengthy screeds from AI and not posting AI responses in FQ,

Or Hamlet, either.

Great summary. Thank you!

I’m wondering if / when the OP will return to grace us with their feedback on our collective comments.

They posted a couple of hours ago, just not here.

To be clear, when I say that I think it plausible that an AI might have inadvertently memorized some of Shakespeare’s plays in their entirety, I’m not expecting absolutely perfect accuracy, any more than I’d expect it of a human who’d memorized one of the plays. And I wouldn’t expect to see it of pretty much anything other than Shakespeare.

I’ve read all the comments.

And do you still believe that LLMs are just storing many zettabytes of copyrighted content and just mindlessly spitting it out?

Because it should be clear by now that this isn’t even remotely how they work.

If anyone wants to deep-dive, there are quite a few papers on this topic. I was looking for data on probabilities of completing Shakespearean works given a seed, but found that Harry Potter is more ubiquitous (at least for llama3.1-70b).

From the Abstract of Extracting memorized pieces of (copyrighted) books from open-weight language models:

This isn’t to say that llama3.1-70b memorized all books – it didn’t. Simply that it fairly well memorized the first Harry Potter book (seemingly multiple editions of the first book).

I’m just reading this paper now and going off the abstract for this post.


Here are a papers I found on this topic. I’m hiding them since the list itself is from ChatGPT:

Papers
  • Rethinking LLM Memorization through the Lens of Adversarial Compression — Schwarzschild et al., 2024
  • Memories Retrieved from Many Paths: A Multi-Prefix Framework for Robust Detection of Training Data Leakage in Large Language Models — Dang et al., 2025
  • Measuring Memorization in Language Models via Probabilistic Extraction — Hayes et al., NAACL 2025
  • Extracting Memorized Pieces of (Copyrighted) Books from Open-Weight Language Models — Cooper and Grimmelmann et al., 2025
  • Quantifying Memorization Across Neural Language Models — Carlini et al., ICLR 2023*

I read the paper, and it’s very interesting and should contain surprises for both sides of the legal argument about copyright. Unfortunately this isn’t the right forum for an extensive discussion, but I’ll just say that a key point the authors make is that both the plaintiffs on behalf of copyright holders and the defendants on behalf of LLM developers are greatly oversimplifying a complex and nuanced situation.

It’s absolutely not the case that LLMs in any sense “store copies” of their training material. What’s astonishing, though, is the extent to which a few materials to which they’ve been frequently exposed can be regenerated from the parametric weightings in their neural nets with reasonable if imperfect fidelity.

This seems to me to be similar to the way human memory works. As a big fan of Fawlty Towers, I’ve watched it so many times that I can recite much of the dialog, but not perfectly, and certainly not as if I’m playing back a recording or reading from a copy of the script. I’m regenerating it from remembered patterns. That’s just what LLMs do, with equivalent imperfections, though it bears repeating that in the vast majority of cases they can’t even do that without reaching out to the internet, because they don’t remember much beyond a vague outline, which, again, is much like human memory.

The massive piles of digital books used in training aren’t necessary closely curated. There’s a strong chance that there were several copies of the book with different file names. Which would influence the “memorization”: if you have five copies of the same sequence of words (or an extremly similar sequence, such as some using the word “Sorcerer’s” and some the word “Philosopher’s”) the LLM is going to weight that sequence more strongly than if it has only one.

Yeah; all that should matter for this forum is that AI results are not, at this time, considered a copyright violation in our legal system. That’s all that should matter for this thread subject.

Whether or not they should be, that’s an interesting discussion that seems (IMHO) outside of the scope of this thread.

That’s highly surprising, given both that any Harry Potter book is much longer than any Shakespeare play, and that Shakespeare is quoted so extensively in basically the entire English language, while Harry Potter, being both more recent and more niche, isn’t nearly so widely quoted.

One can think of it as compressed copies of the training material, but with highly lossy compression, and some originals have more (or less) loss than others. In most cases, the amount of loss is easily enough to make it not a copyright violation, with a very small number of exceptions.

Shakespeare fans touch grass.

And the Bible is even more highly quoted- but not copyrighted (except in a few new translations, etc).

This entire message board is human splooge.

This message board protects direct copying of someone else’s content. You can paste a little bit (fair use with attribution) but that’s it.

AI is doing the same thing. It hoovers up content, same as you would reading a book or article, and then outputs some synthesis of that.

This is EXACTLY what every person on the SDMB has ever done. Exactly what every student in school writing a paper on a topic has done for ages. Why is AI different? Because it can consume a library of books in a few seconds that is now bad? That’s just more efficient. Some people read faster than others. Are they more bad than the slow readers? Where do you draw a line?

If you catch someone copy/pasting someone else’s content let us/the world know.