So... How does one actually use AI?

Lovely forms of women sculped Junonian. Immortal lovely. And we stuffing food in one hole and out behind: food, chyle, blood, dung, earth, food: have to feed it like stoking an engine. They have no. Never looked. I’ll look today. Keeper won’t see. Bend down let something fall see if she.

Butt-Head: He said “hole”. Heh heh; heh heh.

P.S. to be a really useful A.I. search, if should not matter if the sentence specifically contains the word “butt” or “hole” per se, it should just know what is being described…

There are plenty of secondary sources that talk about Leopold Bloom looking up the backs of the statues to determine if they have buttholes. It doesn’t need to evaluate the primary text.

Speaking of AI and buttholes…

(Sorry if the AI logo/butthole connection was already brought up in the thread and I missed it.)

I have to say that ChatGPT and Claude continue to impress me. Here’s today’s example.

I just had a new vacuum cleaner delivered, along with a package of replaceable bags. The model number of the vacuum wasn’t among the long list of compatible model numbers on the bag package. When I tried to check compatibility on Amazon, it returned “cannot confirm compatibility”. Googling was of no help.

So off I went to ChatGPT, gave it the brand and SKU of the vacuum and the bag package, and asked it to check.

What impressed me wasn’t just that it came back and said “yes, they’re compatible”. It was that it returned an analysis of how this manufacturer’s part numbers have morphed over time, and how they’re organized into compatible families, and how the new part numbers and family organization aren’t always reflected in the current documentation or packaging, and finally, how the bag part number directly relates to the compatible vacuum model even if the package doesn’t explicitly say so.

Otherwise I’d have been needlessly schlepping all over town dropping off an Amazon return and ordering a new package of bags that may well have been inferior.

For reasons, I wanted to find an adjective describing chickens, analogous to ovine for sheep or bovine for cattle. So I typed into the Google search engine

sheep:ovine=chicken:?

and got the immediate answer that you could use avian, but the formal term was galine. I was impressed.

I got a 3D printer recently, and started experimenting with CAD. One of the things I wanted to try to build was a small catamaran “beer boat” that I could tow behind my kayak, instead of one of those flimsy beer coolers.

Claude modeled this on its own:

And then ran some computational fluid dynamics on it to model how it’d perform in a river:

Then it kept refining the design… until I ran out of usage credits again :sweat_smile:

It also taught me how to use both tools, connected to Fusion on its own (through the “MCP”) and was able to manipulate the workspace for me to show/hide certain things and walk me through the features. This is a project that would’ve otherwise taken me months or years to make; instead, we got that far in an afternoon.

My understanding is that Claude is actually not very good at this (3D modeling), and that ChatGPT is supposed to be much better, especially their latest model. I’m really tempted to try it, but still bitter about their whole DoD debacle. Meh.


It’s also really good at manipulating 3D models in simpler ways, such as making this sink reinforcement base plate from some photos and measurements (to strengthen my flimsy metal sink for a heavy faucet):

Or creating this calibration test plate for the different “fuzzy skin” settings on my printer:

It’s interesting to me that Claude can do this despite not having a built-in vision generation model like ChatGPT and Gemini have. It can’t make images, but it can still reason about how parameters and numbers in code would look in 3D — much more so than I can in my head, at least.

I eat dinner out with friends every Sunday. We’re in a restaurant rut. I’m having ChatGPT make a database of every sit-down restaurant in the city with type of cuisine, price range for an entree, parking availability, bus route, neighborhood, noise level, open Sunday evening or not, reservations, etc. It’s taking a long time because it’s using several different city and county databases to figure out what are all the food-service establishments, which places are restaurants with sit-down dining, and then gathering the relevant info from their websites. It appears there are 1,282 restaurants it has to gather info for.

I am not. Chicken (or hen) is “gaina” in my language. It obviously comes from Latin.

If you ask AI to devise ten various analogies with different logical relationships between the terms, you will see it stumble.

I doubt ordinary AI programs can do their job accurately when it comes to assessing things about the real world (i.e. space, time, locations, coordinating things in motion, etc.). My wife runs an unofficial tourist agency for our family and friends. She is by nature (and due to her profession as well) a highly analytical person and she gets frustrated by both Copilot and Gemini because they fail to include relevant information or to use it consistently. If you use these to schedule trips and accommodations across the world, you may ruin people’s experiences.

I did an experiment in this respect. I asked AI to count the number of restaurants in a strictly designated area (my neighborhood), between two avenues on one axis, and two subway stations on the other. It is mainy a residential area. They found between 150 and 200 restaurants, which is way too much. They must have added take away places, bakeries, and so on.

This is an interesting experiment. Can you think of an example to try? I would think that analogies would be one of its strong suits (given my admittedly limited knowledge of how transformers and vector space embeddings work).

This is another class of problem altogether, and one that I could easily see LLMs stumbling on, especially if they were limited to a sandboxed chatbot (and not have access to GIS tools and POI databases — purpose-built geographic software and points of interest listings to validate against). And yeah, the definition of a “restaurant” could alter the result set too, with or without LLMs.

That’s what I thought too but they’re not. Or maybe there are more advanced models that I haven’t used. I refer to Gemini and Copilot. When they devise their analogies, they should use various logical relationships, such as synonyms, antonyms, cause and effect, part and while, worker and tool, etc. Try to ask them to work with a given set of words and you will see they end up using only synonyms while the problems they give are often lame or even wrong. Disappointing.

ETA: One lame example is when you start from, say, “serene” to get “serene,” the relationship being IDENTITY. Sure.

What do you mean by that? Like are you asking it to come up with analogies from scratch, or are you feeding it a list of words (from where?) and asking it to determine the relationships between them?

I don’t get it… :sweat_smile:

Here’s Claude (Fable 5.1) attempts at analogies, for what it’s worth:

Seemed decent enough? Probably a better list than I could’ve come up with. It’s actually kinda a fun game (about halfway through the transcript), having it make quizzes for you to answer.

Edit: I also tried it with Haiku, its weakest model. Performed much worse, borderline nonsensical sometimes:

(Claude is the brand; the specific models have different names within that brand. Fable is the strongest/slowest/most expensive, Haiku is the cheapest/fastest/worst.)

I might try that with my town.

We have two restaurants.

So far this year, I haven’t been to either one of them.

The difficulty resides in making them use the words you provide, but I ran experiment now and I must admit that Gemini did a good job.

I did that (later on in the transcript, sorry). Excerpt:

I asked:

We have a doubter – a skeptic, you might say – about your ability to do this. They said:

[snippets of your posts]

So… what about you come up with ten more quiz categories, but I will give you some “seed” words, and from those initial words, please devise a category for it, come up with its related word, and provide the third word but leave the fourth one blank for the human.

Words:

blimp
death ray
chicken
tomato
Rome
cryptocurrency
naysayer
Coptic
black hole
rangefinder

Its list:

Ten, each a different relationship, none of them synonyms.

  1. Blimp : buoyancy :: airplane : ____

  2. Death ray : mad scientist :: grappling hook : ____

  3. Chicken : cowardice :: mule : ____

  4. Tomato : fruit :: peanut : ____

  5. Rome : Carthage :: Athens : ____

  6. Cryptocurrency : mining :: fiat currency : ____

  7. Naysayer : proposal :: heckler : ____

  8. Coptic : Egypt:: Aramaic : ____

  9. Black hole : singularity :: hurricane : ____

  10. Rangefinder : distance :: anemometer : ____

(answers in the transcript)

I’d be curious of the result.

But IMHO, since we’re in a “how to AI” thread… I believe what that is primarily testing isn’t the strength of the LLM model itself, but:

  1. whether this is a famous fact about your city that found its way into the training set (a factoid like how “Paris has one Eiffel Tower” that everybody knows)
  2. how good the “harness” is, the software around the LLM that helps it with complex information retrieval & analysis, by providing it non-LLM classical software tools that it can access (like a mapping utility or a database) — or not
  3. whether the particular AI brand/company you use wanted to provide such a capability — which is mostly a business and cost decision, because those workflows cost more money than something the model can just pull out of its training without external tooling

I think this is the sort of question that would be answerable relatively definitively with a harness like Claude Code or OpenAI Codex, but very difficult for a bare chatbot of any provider (except maybe Gemini? Google is building stronger integration between its AI and Google Maps; I haven’t tried it). But generally speaking, the chatbots don’t and can’t build themselves the tooling they need to properly investigate a question like this.

It’d be like locking the world’s best detective inside a closed room with no access any outside information, and asking him how many restaurants are in a city he’s never been to or heard of. LLMs can’t magically produce an answer for that out of nowhere either, without the harness giving them a way to look that up directly, or create for itself a way (like a small program) to do that lookup.

With an agentic coding harness (instead of a chatbot one), any recent, half-decent LLM would be able to write a geospatial lookup tool that defines a boundary polygon precisely out of government-provided census shapefiles that define a township (or some arbitrary boolean intersection of roads), combine that with a POI listing (or OpenStreetMap data), and do the right calculation to give you a correct table (assuming a sufficiently tight definition of “restaurant”, and a good enough source dataset).

You don’t need a particularly strong model for this, and even the weaker ones know how to do something like this, they just can’t because the typical chatbot experience essentially traps them in that locked room with very limited access to the outside. They can do simple web searches (which return results by relevant/SEO spamminess), but they can’t do a proper geospatial lookup like this because they are artificially constrained. That same model working in an appropriate harness will have much more freedom (and likely higher success rates) for something like this.

It’s the difference between:

  • “Do a web search and try to find all the restaurants in my city” (the chatbot experience). I don’t think a human would succeed at this either, because it’s difficult/impossible to list ALL the restaurants from a search alone. It is an artificial constraint imposed by the chatbot experience, not because the AI is incapable.
  • “Solve this same problem, using all the tools known to you, and you can write your own software to help you” (an agentic harness). It’s no longer limited to a web search and can do anything it needs, and pull data from anywhere it wants.

So all that is a very long-winded way to say that the AI providers provide not only different models, but also different harnesses to use to tackle different kinds of problems, but they don’t really guide you through the differences anywhere. You basically have to be a software developer to know that these other harnesses even exist, and when would be appropriate to use them vs a chatbot. These are productization failures, not necessarily LLM ones…

Well, that’s very bad because that’s how Copilot is advertised, for example. I use Microsoft Edge and I’m satisfied with it, but there are frequent updates. Every time it gets updated, we learn “what’s new.” Copilot’s enhanced capabilities are always “new.” I haven’t memorized what it says, but they claim we can use Copilot to organize your work, your day, your vacation, everything. Ask the Copilot the fastest way to get to Times Square from JFK Airport, and then compare its answer to what a new-yorker you can trust will tell you. You may be surprised.

Yeah, I know :frowning: Agreed 100%.

Microsoft is probably the worst of the bunch, honestly, a leaderless company that’s been adrift and directionless for the last several tech revolutions. They mostly just invest in other companies and repackage their models.

But its not just them. Nobody really knows how to sell AI effectively and profitably and sustainably yet, so they’re just kinda throwing random workflows out right now and seeing what sticks. But it’s 90% marketing fluff while the actual useful details are buried in that last, barely mentioned 10%. The actual AI capabilities today far exceed the average person’s experience with them, because the providers haven’t quite figured out how to effectively advertise and teach the more advanced stuff yet. The best is all locked behind special tooling and pricier (but still extremely subsidized) plans.

If Cecil were still around, this would make for a very good topic for a SD article :frowning:

I guess the question is, does everyone need/want its more advanced capabilities?

In the arena of average users, I used Claude today to help me figure out a system for monitoring active MDHHS grants. This was not particularly sophisticated - I did it in Trello, but now I can see at a glance not only which grants are active but what phase they are in and whether there are any dependencies. The main reason I did this is that these grants are always going through different phases and usually it’s someone else who is responsible for each phase, so I needed a way to make sure I can get on these people’s cases if there is anything outstanding. And what is technology useful for if not harassing your coworkers?