Is it possible to determine who is responsible for the largest number of replies across all threads?

We have a number of threads that have exceeded the 10,000-post limit and gone on to Parts 2, 3, 4, and more. IMHO, even if the OP of the initial thread didn’t actively start the follow-on threads, they are responsible for all the posts in all the threads that followed.

So the question that comes to mind (my mind, at least, such as it is) is, which Doper has created threads that have the highest combined total post count?

Is there a way that could be calculated and determined?

At a quick glance through the tools I have available, I don’t see anything that would easily work. I’d have to probably do it manually, the same way you would - track the 10k plus threads (using that as a rule-of-thumb cutoff), find the op, jot them down, repeat. Which in a way could be a reasonably fun thing to do and research.

If you’re really curious (and the mods are OK with this), I could write a script to do the manual work that ParallelLines would’ve done by hand — unless they want to do it by hand :sweat_smile:

It wouldn’t need any special access, just start from something like the top topics (or the machine readable JSON version of the same), look up each one’s OP and tabulate the replies, and then paginate through the topics list until we’re statistically pretty certain that probably there isn’t a long tail (of like some poster who miraculously made a few thousand topics with only a couple replies each). (Edit: Actually, Claude noticed that https://boards.straightdope.com/about.json says we have about 28 replies per topic on average, so the long tail DOES matter here and shouldn’t be ignored.)

It would be a “polite”, well-behaved crawler that shouldn’t cause any noticeable load on the server, and it could sample pages until we build up a statistical confidence of our results, maybe stopping short of a full crawl. A full crawl wouldn’t be that bad anyway (since we’re fetching metadata of topics, not all the replies within it) and would take maybe 2-3 hours total to do — mostly because we’re just trying to be polite to the server.

Let me know if you want me to build this.

Also… if we actually had a board admin that was still active (i.e. a user with SQL access to the Data Explorer Discourse plugin), it would take a few seconds instead of a few hours… sadly I don’t think we do :frowning:

I have the least experience with the tools, so there may be a better way to do it. I certainly don’t have the time or energy to do it justice, well, maybe at the end of the month when I’m doing a dual Flu/Covid vax, but I’m normally quite under the weather after that.

I suspect that there’d be enough privacy/liability concerns with even the most benevolent crawler for TPTB to be extra cautious with such, absent a greater need than our (unquenchable) curiosity.

To be clear, the crawler wouldn’t access any private data at all. It wouldn’t even be logged in, much less have any sort of special permission. It’s just looking at the topics lists that any other person or bot could see, malicious or otherwise — I’m sure far worse have already tried, and we’re already all in AI training sets, for better or worse =/

But, still… I made a similar, presumptuous (and arrogant, yeesh, I was such an ass back then :sweat_smile:) dismissal when I was younger here: A browser addon that embeds SDMB user portraits into threads?

I won’t make that mistake again. People can get nervous about their privacy, so if there’s a concern about it, probably better to just respect them.

Besides… it’s all about quality, not quantity, right? :slight_smile:

He says, pointing to a post from 2010, while I have only been posting since 2020. And meanwhile the 99’ers look down upon us from the lofty heights of epic senority!

Seriously, I absolutely believe your best intentions and question them not at all. Just that we tend to be very cautious, exactly as you suggest.

Well, if the knee PT earlier today didn’t do it, now I certainly feel old :wink: Though I’m not really; I just started posting here as a teenager, and a lurker for even longer before then. It’s been a while!

If the SDMB (and internet at large) has taught me anything, it’s that even the best of intentions are rarely read as such online. People will often default to the least charitable interpretation. Something about the facelessness of pseudonymous texts on a message board, I guess?:person_shrugging:

Anyway, the OP’s question is interesting enough in its own right, but not worth a possible uproar over. Maybe someday when we’re all dead, our kids and AIs will make a SDMB museum where this can all be calculated and preserved as a part of early internet lore…