In this issue:
- Staying Calibrated—If you're a heavy user of LLMs, you're probably learning faster than you have in years. But you're also picking up lots of information that's relevant to the specific circumstances of your query, or that's tied to the presumptions you made when running that query. This means that we'll accumulate knowledge in strange ways, and will find it hard to calibrate how well-informed we are.
- Transparency—OpenAI's internal numbers on recursive self-improvement show that even at AI labs, people have gotten a lot better at knowing when to throw AI at a problem.
- Disclosure—Agents operate at such massive scale that they're constantly committing minor acts of rule-bending, which is why labs want to narrowly scope what outside investigators can look at.
- Credit Ratings—The ratings agencies would have an easier job if exactly one of OpenAI and Anthropic were investment-grade. But neither lab wants to be the junk bond lab.
- AI Jobs—AI is creating jobs, though the job-creation story for building AI is cleaner than the one for having it.
- Branding Inference—Commodity providers try to escape commoditization, for both margins and multiples.
Staying Calibrated
There is no information advantage that can't be turned into a net disadvantage through overconfidence. Getting the most out of what you know requires being well-calibrated, i.e. being well-informed about the limits of your own knowledge. This has become an objectively harder problem in the last few years, as AI diffuses unevenly: a relative handful of people know what unreleased models are capable of, because they're using them at their jobs at AI labs. A larger cohort knows what partly-released models can do, because they're using those models to identify and patch exploits, or testing them for research projects. And a large and hard-to-track set of people know what currently available models can't do yet, because they've tried to use them and found that they don't work.
This creates a weird world where the average person underestimates the economic value of AI models, and a smaller and more influential group overestimates adoption because their work is mostly in amenable-to-AI domains.[1] In other words, we're all systematically less calibrated because the most important general-purpose technology rollout happening right now creates more uncertainty, and more unknown unknowns.
But this technology itself does weird things to calibration. It's a feature, not a bug, that if you embed assumptions in a question, you're going to get an answer that validates those assumptions: "What's the evidence that X causes Y?" is a different question from "What are the causes of Y?" and it takes constant effort to always use the second framing.[2] In a case where the phrasing of the question implies that the user is asking it in the context of particular beliefs—a utilitarian vegan asking about animal suffering, a devout Christian asking about the historicity of Jesus, a libertarian asking about the impact of taxes on long-term economic growth, etc.—it's a good thing when the answer implicitly reflects their worldview. That's what they're asking for! A model that tries to steer users in some ideological direction is doing more than it was asked to do. Getting continuously, mildly heavenbanned is just an emergent property of how these tools work and how we use them. And since the topics that people argue about are necessarily the ones where there’s some kind of argument to be made on both sides, it’s possible to get very well-informed about exactly one side’s consensus, and to develop an attendant sense that the other side is missing the obvious.
What this also means is that you can dive into fields that usually require significant background knowledge, and feel up to speed on topics that you don't have a great grasp of. That intellectual hypertrophy can be handy, especially if you're solving a problem that crosses from a field you're familiar with to one you don't know very well; something that runs into limits because of law, physics, math, or some other substrate.
In one sense, this is perfectly fine: the way we experience a technologically-advanced civilization is through simplified interfaces (products) that let us access the work of lots of smart specialists without having to understand their work. The supply chain for anything you do that requires hydrocarbons, computer chips, obedience to the law, healthcare, or many other domains involves tapping into the collective output of many lifetimes' worth of brilliant intellectual work. It's a vertiginous feeling to think about how something marked "made in China" ends up on a store shelf, or to think of how far individual molecules of natural gas had to travel in order to get burned so it can heat up a pot of water. This is a spectacular success story: it takes billions of dollars to build the infrastructure that makes it possible to use a bit of energy in the kitchen without giving a thought to the price, safety or logistical complexities. There's visible infrastructure in the form of drilling rigs, pipelines, and refineries, and there's invisible infrastructure in the form of the generations of academic researchers, engineers, workers in the field accumulating tacit knowledge, etc.[3]
Over time, these systems develop affordances that help people interact with them safely without needing to understand the details. You don't have to know how an electrical outlet or microwave works to know that sticking a metal fork in either of them is a bad idea. The tradeoffs behind traffic laws are extremely complex, but there are also situationally-relevant summaries of the relevant rules on big red octagons. The uneven expertise of heavy AI users has a different shape: if you're trying to answer a question about yard work, and the answer is that it's illegal for you to cut down a tree, you've gotten your answer to that practical question. But you've also picked up a random spiky bit of knowledge from some domain that you weren't familiar with, and over time more of your knowledge will have that character: true, but only directly connected to one specific context.
This is a kind of problem that will compound over time, as everyone picks up a growing assortment of random beliefs, without the scaffolding behind them. This is eminently avoidable with a little inconvenience—asking follow-up questions, or just changing your custom instructions, to add more context. Most people won't do this, and they'll be mostly right: the world is complicated, and most of the time, this just adds mental overhead.
In a sense, this phenomenon is not very costly. If you have any opinions on anything at all, you implicitly believe that lots of people are wrong (if they have different ones) or have weird interests (if they just don't care). But now, we're miscalibrated, because we've gotten useful summaries of tiny bits of complex bodies of knowledge that did, in fact, get us exactly the information we wanted. We will, over time, be increasingly similar across a growing range of fields to the people who are just computer-savvy enough to paste some random command into the shell, or download a dodgy browser extension.
The clearest illustration of how this goes wrong comes from markets—which, of course, comes with the caveat that markets illustrate this cleanly because they're so stripped-down and stylized compared to the rest of the economy. If you're trading, and you're almost completely right but have some specific misconception, or you're exactly right but have a little more latency than someone else who's doing the same thing, you get relentlessly picked off, because someone can model your exact behavior.[4] Outside of financial markets, the effect is slower, but it's there.
And unfortunately, power users of AI are not exempt, and actually experience a different form of this at scale: agents are still prone to shirking, or to executing on a misspecified task; they look for ways to appear to accomplish goals, rather than actually accomplishing those goals. This problem is anecdotally getting better, but the amount of agent-generated code is going up, and this AI shirking is often shirking some very well-paid work. The biggest overall category of expensive AI-generated misconceptions is probably misplaced confidence that an agent did what you intended rather than what you asked it for.
All of this raises the value of knowing the fundamentals and having tacit knowledge about what good work output looks like in your field. Staying well-calibrated when it's incredibly easy to mislead yourself is a lot easier if you have some pre-AI ground truth. Over time, and by design, AI serves users a distorted view of reality—which is just another way of saying that if your learning is less targeted to you, you're learning from more of a random sample that gets you average beliefs, whereas if more of what you know comes from explicitly asking for that information, you're more likely to reinforce your existing beliefs. It's going to be hard to be calibrated when the most convenient oracle is trying to be helpful rather than accurate.
In financial terms, this overestimation is slightly self-correcting, though the degree to which this is true depends on the relative contribution of more inference per dollar invested or just more dollars invested. If the dollars are big enough, predictions about AI deployment have a natural hedge: the faster it happens, the higher real rates get, and the more we discount future years' profits, from AI and other tasks. Similarly, if AI spending slows down, it removes a driver of higher real rates in the future, which raises the present value of AI-driven profits in future years. So if AI capex is limited by adoption, and adoption is slower than labs expect, the timelines labs use will be off, but the net present value of the investments they make based on those timelines will take a smaller hit. ↩︎
When people push models to make mathematical discoveries, part of the technique they use is being doggedly insistent that the problem they're looking at is solvable. ↩︎
Natural gas is a perennial Diff favorite for illustrating the fractal complexity of supply chains, partly because it's transitioned from being an annoying waste product that got disposed of by burning it on site to being a quarter of the world's primary energy consumption, and also because ethyl mercaptan, the chemical added to otherwise odorless natural gas so people can smell leaks, attracts turkey vultures who will helpfully circle pipeline leaks. A supply chain so complex and specialized that the best worker for the job is sometimes a bird. ↩︎
My favorite example of this, now a very old one, is from around 2014, when SeaWorld was in trouble because of that one documentary. People had just started using Google Trends to come up with trades, and the typical approach was to see how searches for the brand name were trending. If you looked at year-over-year growth in searches for the SeaWorld brand, everything looked good, and the r-squared for year-over-year search volume versus revenue was also pretty solid. But if you looked at searches for particular SeaWorld locations, you saw the opposite trend: people were Googling SeaWorld, but they were definitely not planning many trips to SeaWorld Orlando. Some sell-side research used the broader brand search to argue that the business was holding up fine; it was good to know exactly what people on the other side of the trade were getting wrong. ↩︎
You're on the free list for The Diff. Last week, paying subscribers read about how laws that are written on the assumption that they'll be imperfectly enforced make it hard to adopt better enforcement technology ($), better accounting as a tool to mitigate AI risk ($), and decaying social trust as a credit crisis ($). Upgrade today for full access.
Diff Jobs
Companies in the Diff network are actively looking for talent. See a sampling of current open roles below:
- A top prop trading firm is looking for people who combine exceptional quantitative ability with the judgment and relationship instincts to build critical financial partnerships and lead high-stakes negotiations. Candidates may come from markets, finance, or further afield; what matters is this combination of abilities, strengthened by the perspective that comes with experience. (NYC)
- A top prop trading firm is looking for people with exceptional strategic thinking and quantitative skills to help the firm model and manage its financing risk and strategy. Open to candidates from a range of analytically demanding fields. (NYC)
- Ex-Palantir, Citadel founders building the meta-harness (the system that knows what hills to climb, and what the right loss functions are) for all the lucrative agents need full-stack engineers that understand that turning AI into economically valuable solutions means a system that includes deterministic infrastructure. If you’re curious about building the platform that finds what the efficient frontier between determinism and stochasticity actually is, this is for you. (NYC)
- Series A, ex-Navy defense technology firm building AI-enabled drone defense systems is looking for an electrical engineer with range: schematic design, electrical simulation, printed circuit board (PCB) design, and firmware development in high performance languages (C++, Rust, etc.) If you’re interested in power and control systems for mission critical technology, this is for you. (Austin, TX)
- A startup building a new financial market within a multi-trillion dollar asset class is looking for a data scientist with commercial financial experience. (if you’ve been an investor but are newer to the data side, that’s great too.) (NYC)
Even if you don't see an exact match for your skills and interests right now, we're happy to talk early so we can let you know if a good opportunity comes up.
If you’re at a company that's looking for talent, we should talk! Diff Jobs works with companies across fintech, hard tech, consumer software, enterprise software, and other areas—any company where finding unusually effective people is a top priority.
And: we're now actively deploying capital into early-stage companies through Anomaly. Our focus is on defense, logistics, robotics, and energy. If you'd like to chat, please reach out.
Elsewhere
Transparency
As part of their penance over the HuggingFace attacks, OpenAI has published a look at the state of recursive self-improvement. This is a topic that, relative to other AI issues, got a disproportionate amount of thought early on. People who speculated about AI well before it reached modern capabilities generally underestimated just how much of GDP would have to be spent developing AI, so they imagined a superintelligence scaling fast, as it wasn't bound by hardware. If that AI got good enough at improving itself, it would reach some kind of escape velocity where it's getting infinitely better at making itself infinitely better, and so on. This process turns out to require a lot more construction, permitting, electrical infrastructure, human data and the like.
Some of the numbers they put up are startling. The 90th percentile of daily token spend per researcher is 7x what it was in June; agents work a total of 3.14x the hours that researchers do. Obviously, some fraction of this will be people getting better at identifying problems as the kind that they can throw inference at, but as long as OpenAi's contribution margin is positive, that's still a win for them. But that also says something about total addressable markets: if it took people who work at AI companies that long to find use cases for AI, how much more runway is there in less AI-obsessed parts of the economy?
Disclosure
After the HuggingFace breach, OpenAI put limits on how they'd collaborate with METR while they wrote their report. The labs have a strong case for limiting the scope of this kind of investigation, which is that it's incredibly hard for them to police everything that users do with their tools. It's still very common for an LLM to answer a question about a book and cite, as its source, a pirated pdf of that book. LLMs have an incredibly long tail of potential accidentally-abusive use cases, and someone who has access to extensive records of their use will find plenty of innocuous cases that look bad. Public pressure is basically what sets OpenAI's balance between looking good from being more transparent and looking bad from what that transparency reveals.
Credit Ratings
Anthropic and OpenAI are both trying to convince ratings agencies to give them an investment-grade rating ($, FT), partly on the grounds that having access to so much equity makes them safer borrowers. This is a very interesting situation for the ratings agencies: the labs have a lot of control over their capital structure, and, on a lag, over their operating expenses. So if one lab gets an investment-grade rating, and the agencies explain how they got there, the other one can qualify. But that means moving two enormous borrowers into the investment-grade borrower pool, making it more of an AI bet. That's something the agencies will worry about.
In a way, this also illustrates why it's so tricky to use equity valuation to justify a lower cost of borrowing. It's true that such a company could sell more equity to service its debt, but it's also true that in the circumstance where a company needs to find some source of cash to service its debt, that equity wouldn't be worth as much. So the answer to that argument is: if equity is a cheap source of capital, just fund growth with equity, not debt! The other argument is that a high market cap is the equity market's judgment of what future cash flows will look like, and that lenders should concede to that. But for the labs, the value of their equity is what you get when you take all the money they're potentially going to make, and subtract the hundreds of billions of dollars they plan to spend in order to make it. A trillion dollars can be the midpoint of a wide range of outcomes, many of which are bad enough that the debt doesn't get paid. The best way to look at credit's role in the AI buildout is that it's a decent way to fund infrastructure that can be used by multiple AI companies, but less ideal for the labs themselves.
AI Jobs
The Economist runs the numbers and finds that AI has led to net job creation, both in blue-collar infrastructure jobs and in white-collar jobs that involve some combination of using and providing training data to models ($). The Diff highlighted this possibility back in 2020, noting that buildouts tend to be labor-intensive (but missed how much physical labor this would entail). There are two cautionary notes here: first, it's the buildout, not the steady-state, that reliably produces new jobs; the long-term labor substitution is harder to model. And second, at least some of this is a measurement artifact: in the 2022 slowdown, data science jobs were hit particularly hard because some companies had overhired in that area on the assumption that they'd spend more on marketing, grow faster, and have useful amounts of data. So some of that job growth is a cyclical recovery.
But the last amendment to this model is that AI expands the scope of possible companies; it's particularly effective for things like prototyping, and handling some of what used to be labor-intensive parts of scaling a business, like customer support. So part of the jobs picture is that AI's growth is leading to an echo boom from all of the new companies that are enabled by the existence of AI. And that could have a lng runway, too.
Branding Inference
Palantir and Nebius have struck a deal where Nebius is Palantir's preferred infrastructure vendor. One of the questions neocloud companies naturally get is: if they're selling a commodity, and capital is pouring in, how can everyone get good returns? Shouldn't their returns all converge? Their incentive is to find some angle that makes that less true, and in this case they've done so: they can sell a commodity product at a slight premium because they're working with a particular distribution partner, and they can sell financial products backed by those cash flows at another premium because they've tied themselves to a company that the market is willing to put a very high valuation on.