Subject inspired by discussion in several threads about AI but not necessarily about AI. It also is a question about how cognitive work productivity is measured, how it actually makes sense measure it, and how productivity and compensation interact, or do not.
Some is just factual questioning - getting some with the chops of understanding economics to clearly explain how productivity is measured for cognitive services. And then, since that is likely GDP based, if that adequately captures productivity, or if it shortchanges outcomes and quality of life improvements gained via the service. Lastly of course how AI as a tool may impact productivity and how to measure it beyond GDP. Assuming it doesn’t go rogue and pursue its own goals.
My specifics are illustrated with medicine just because it is what I know but interested in the more general case.
Consider primary care physicians. Productivity is often measured by Relative Value Units of the work performed (RVUs): a certain service, be it a procedure, a preventative care service, a level of complexity of illness visit code, is granted a certain amount of work value, and within each specialty we are often compensated at least partially based on our productivity as so measured. Primary care docs are doing more in these well care visits coordinating care and various preventative metrics than in the past. But the throughput doesn’t change much.
Is the economic value of primary care best measured by how much the services are compensated? visit number? diagnoses addressed? appropriate care provided? avoided hospitalizations? quality-adjusted life-years? reassurance and counseling? coordination with specialists? prevention of future disease? …
How do we best measure how AI impacts productivity in something like primary care? And by extension other cognitive work fields?
These are what we call in the grants business process objectives. Sometimes that’s all we’ve got, but the real test of an intervention is demonstrated change as a result of said intervention. I haven’t a clue how this applies to medical practice. Also I recognize there’s a difference between productivity and effectiveness. So maybe we’re using the wrong metric?
Cal Newport, productivity guru and computer scientist, has written about the challenges facing knowledge workers and productivity measures for a long time. If you work in manufacturing, you make a widget. That’s productivity.
A lot of knowledge work jobs are process oriented, so bosses, at a loss for any other type of measure, have turned to surveillance methods, email responsiveness etc. In other words productivity theater. We can see an expansion of this in how some companies are measuring productivity by #s of AI tokens use.
Cal’s take on this is what we need to define our own measures of success. For an academic, this might be papers published/articles written. For me it might be applications submitted. My boss is measuring me by grants revenue, but that’s more like effectiveness than productivity. This being a numbers game, the more applications the better. So my goal is just to get stuff out there.
I dunno how AI fits into this. It certainly helps with my productivity, but in a lot of ways I might define as soft skills. It helps me organize projects, does a decent job reviewing narratives, and cuts through a lot of the tedious spreadsheet and flowcharts part of my job. But I’m not getting more done by saving time exactly. I’m getting more done because I’ve created a workflow where I’m not so overwhelmed and avoidant of my work anymore. Claude’s major contribution to my productivity is that it helps me get started and helps me get unstuck when I get stuck.
One issue I ran into at an old job was that we were measured by how many open cases we closed a day. This lead to a situation where everyone was picking the easiest cases and ignoring the more complex cases. My understanding medicine has ways to compensate for this, where more complex and sicker patients give the provider more rewards for managing their health issues, but thats a major flaw in a system that rewards cognitive productivity, people just cherry pick the easiest problems to solve.
This paper talks about how despite the fact that the human capital and financial capital we devote to solving problems keeps going up, the rate of progress itself isn’t growing. The reason is that each layer of problems is more complex than the layer beneath it, and the returns relative to investment are less with each new layer of complexity. We devote around 20-25x more R&D into agricultural crops like corn and soybeans than we did in the 1960s, but the growth of bushels per acre is the same as it was in the 1960s because the problems keep getting more and more complex.
After you solve the easy problems, the problems get harder and the return on investment gets smaller. Then when you solve those problems, the next set of problems is even harder with an even smaller return on investment per unit of human capital and financial capital devoted to solving it.
With primary care you also have issues like patient non-compliance. If the patient doesn’t comply with treatment then they will not show productivity in their health, wouldn’t that incentivize providers to drop patients.
Fair compensation and how to think about what is fair, how to motivate the work that actually delivers returns, is definitely at least related to the question.
But to my greater confusion is on the zoom out from individual productivity and compensation for it, to economy wide?
When we worry that low birth rates and a greying population will require either more workers (immigration) or (?and?) greater productivity per worker, what are really meaning by “productivity” because I don’t think it is really GDP, and it seems a bit harder to grasp when we are producing more services,IP, and both cognitive and emotional “work”. - in that context what does AI increasing productivity mean? Yes it might result in more widgets per capita. More R and D per capita? Better quality of life outcomes per capita?
Any first blush of thinking I understand what it actually is turns out to be a wisp as soon as I look directly at it.
I think that “productivity” really translates to “revenue.”
“We’re making more money, therefore we have become more productive.”
I would question this assumption, but when I see charts tracking worker productivity against real wages, with productivity steadily increasing and wages stagnating, I think they are probably measuring productivity as revenue.
I’m not sure what other metric they could use to track the productivity of all workers everywhere. It would have to be something universal to all jobs.
This article is from 2023, but I feel its relevant.
In total, they visited 17 different doctors over three years. But Alex still had no diagnosis that explained all his symptoms. An exhausted and frustrated Courtney signed up for ChatGPT and began entering his medical information, hoping to find a diagnosis.
“I went line by line of everything that was in his (MRI notes) and plugged it into ChatGPT,” she says. “I put the note in there about … how he wouldn’t sit crisscross applesauce. To me, that was a huge trigger (that) a structural thing could be wrong.”
She eventually found tethered cord syndrome and joined a Facebook group for families of children with it. Their stories sounded like Alex’s. She scheduled an appointment with a new neurosurgeon and told her she suspected Alex had tethered cord syndrome. The doctor looked at his MRI images and knew exactly what was wrong with Alex.
“She said point blank, ‘Here’s occulta spina bifida, and here’s where the spine is tethered,” Courtney says.
Endless thousands of dollars were spent, large numbers of man hours and 17 different doctors all to have chatGPT solve it. Granted chatGPT used some of the medical imaging that was used, but the amount of time and money it took chatGPT to solve this medical problem was far less than the amount of time and money it took all the medical professionals to try to solve it.
Another example is that historically it would take a PhD student their entire graduate career to determine the structure of a single protein. Alphafold determined the structure of 200 million proteins. The amount of time, labor and money it would take to do that the old way would be overwhelming.
In certain areas, AI can drastically reduce the amount of financial capital and human capital required to solve a problem, meaning more problems get solved per the same unit of human capital and financial capital. We used to have about 80% of the public working in agriculture in the late 18th century. By the year 1900 it was about 30% of the public. In modern times its about 1-2% of the public working in agriculture.
At the same time elder care is very labor intensive, and until we have affordable advanced robots, AI that can only perform cognitive work probably won’t be that helpful. However cognitive AI could help us design better robots to care for the elderly, freeing up people to perform other jobs.
There are a few problems with this. First, how do you count activity put into a product that flops, which has little or no revenue? When I was at Intel I worked on a massive project to do a new architecture. People designed the architecture, drew circuits, managed it, wrote tests for it.It eventually got out but created very little revenue. Did all that work count as productivity or not? I’ve known lots of chip and processor designs which got canceled before being produced. Is that work productivity? And you obviously can’t measure it for years if you are measuring revenue.
I think wages tracked productivity more back when we had unions.
And the vignette shared by @Wesley_Clark also illustrates the limits of conceptualizing productivity as revenue produced.
By that metric the multiple doctors producing multiple visits and multiple tests ordered without a correct diagnosis generated substantial revenue so were therefore more productive than the maternal AI report informed neurosurgeon was. The primary care providers who order many unnecessary tests that beget concern and follow ups and needless medications when no testing or medication was actually indicated are more productive than the thoughtful experienced provider who accurately recognizes something as benign and lets things be unless it starts to not follow expectations.
It may be revenue is the means that productivity is defined economically (I defer), but to me that is not actual productivity?
Oh, yes, I agree. I wasn’t sure if you were trying to define productivity in the sense of, “how will macroeconomics capture the value added by AI?” Or in terms of, “How do any of us know we’re being productive in the era of AI?”
Trying to self educate into a rabbit hole on ChatGTP and would hope for some real knowledge to verify and expand:
two concepts solve different parts of the problem:
Quality-adjusted output indices ask: How much more valuable/useful is the medical service being produced?
Total factor productivity (TFP) asks: How much more output are we getting from the total bundle of inputs used to produce it? …
… Statistical agencies use hedonic/quality-adjustment techniques to estimate the effective price and quantity of the improved product.
2. Medical care makes this much harder
For a computer, you can measure:
processor speed
memory
storage
display characteristics, etc.
For a physician, you might want to measure:
diagnostic accuracy
survival
symptom improvement
complications avoided
functional status
patient experience
appropriate preventive care
downstream utilization avoided
That’s why economists increasingly favor outcome-based measures of medical output. …
And it goes on to explain about the Solow residual which is the costs of the inputs to a system, say increased staff or capital, minus the total quality adjusted output productivity.
Among academics, the first measure is publication. Even more important is citations. But there are ways to game this. I once sat for three years on a committee that awarded research grants. Some clown came in with 90 papers over the preceding five years. We send the applications along with five sample papers to external referees. One referee commented, “Couldn’t he find even two papers among the five that were actually different?” We turned him down, he appealed and the pointy-headed bureaucrats overruled us.
I once worked for a large bureaucracy (about 20,000 in all offices) where we had a chief who was famous for saying “what gets measured gets done.” He wasn’t wrong, but he’d also convinced those responsible for hiring him that that was pretty much the only metric that mattered. Unfortunately, there were several units in the organization whose jobs were to think of smart things. It’s hard to turn abstract thought into widgets, so the Thinkers of Smart Things were often called into the carpet for not producing enough widgets. “We thought of these Smart Things this reporting term” was their response at first, but since some of those Smart Things were not products, the Thinkers were often kicked to the curb. Fortunately, that guy didn’t last long, but in the five or six years he occupied the corner office, he did a lot of damage.
One metric for academics that attempts to balance publication quality and quantity is the greatest number n such that you have at least n papers that have each been cited by others n times.
There’s a hidden productivity metric that is very difficult to track in cognitive work. My team and I spend a significant portion of our time doing it, and I’m sure many of you do as well. In my case, we happen to manage an enterprise custom case management system for a huge federal government agency. Lot’s of rules and regs, as you can imagine.
There is a constant churn of ideas that come to our team about how to “improve” this, that, the other. Disproportionately large amounts of our time is spent talking executive leadership and middle management out of ideas that have been submitted in their off-site focus groups and process improvement “ideation” conferences. Senior leadership often doesn’t understand the complexity of the federally mandated processes, let alone the application that we design, build and support. So they come to us with a list of complaints from the so-called community-of-practice focus group (read: squeakiest wheels that just so happen do the least work) and tell us to fix it. 8 times out of 10 it’s not a thing that needs fixing because the functionality already exists and they just don’t know about it or are too work-shy to learn how to use it. The remaining 2 out of 10 falls into two categories: 1) an idea so stupid that it’s hard to even address with a straight face, and 2) yeah, let’s do that.
All that to say, there is significant cognitive work that gets done that helps avoid unnecessary, often expensive, usually wrong-headed work. Nobody measures, AFAIK, the money & time spent not doing something because it’s a bad idea.
Where I have to give credit to AI is the generation of white papers containing word salad that executive leadership won’t read, much less understand. Not even the ExecSum version. That in my view has been AI’s greatest contribution to humanity thus far.
That’s the hard one, isn’t it? There are some jobs, like say medicine, where a great deal of knowledge and expertise is required, but you are for the most part dealing with repeated instances of something which is somewhat known, and the outcome can be measured in various ways.
But design engineering? The whole point is to produce something that hasn’t been done before.
Sometimes an engineer is doing their best work when they appear to be staring into space doing nothing obvious. Managers hate that…
And yes measuring the productivity of that individual is difficult but not as difficult to measure the output of the system that includes the person.
How many novel concepts that made it to various benchmarks including commercial success did the team create at what costs and at what estimated current and future revenue?