Rethinking the Data Moat
I want to highlight a couple of pieces which I found to be quite intriguing. The first is Dwarkesh Patel’s conversation with Ryan Greenblatt, chief scientist at Redwood Research. Admittedly, while the conversation about whether automating AI research triggers recursive self-improvement was thought provoking, it was also quite spooky at times.
The second piece that I would like to highlight is a talk by Shuchao Bi, titled “Advancing the Frontier of Silicon Intelligence: Past, Open Problems, and the Future”. Bi co-founded YouTube Shorts at Google, ran multimodal post-training at OpenAI, and now works at Meta Superintelligence Labs.
While Greenblatt’s conversation with Dwarkesh was published last week, Bi gave the presentation more than a year ago. Since I happened to stumble onto both of these during the weekend, I could notice a healthy dose of similarity in Greenblatt’s and Bi’s arguments. Last month, I wrote about the salience of data in the context of AI and Alphabet bidding for bankruptcy auction for Spirit Airline’s data certainly corroborates to that. However, both Greenblatt and Bi made me re-think my position a bit on this topic. Greenblatt had an interesting thought experiment: if you could hold compute or data constant, how much the model would still improve? That delta of improvement could be labeled as “algorithmic progress” and he made the case that it is a very important driver of AI progress over the last few years (emphasis mine):
“GPT-3 was released in 2020, so it was trained about six and a half, seven years ago. It’s worth noting that GPT-3 is maybe a little too far in the past, but let’s go with this for a second.
If we were to train a model with GPT-3-level compute today, how good would that model be? My understanding, based on how algorithmic progress works, is that we’d be able to train a model that’s as good as the best model we had perhaps around three years ago. So I think that right now we’d be able to train a version of GPT-3 that’s probably somewhat better than GPT-4, a moderate amount better than GPT-4. I think that’s about right. That roughly lines up with how algorithmic progress has worked.”
Bi didn’t quite say “algorithmic progress”, but pointed out that the raw data is “unlikely to be the best data distribution”. He suggested that the incremental improvement in scaling law may come from changing the data distributions or to say it differently, “by equalizing intelligence per token”. My read is that they are essentially alluding to the same argument but using different words to explain their intuitions.

Later in the conversation, Greenblatt expanded why he is not a big believer of the role of “human expert data” in model improvements, rather the process improvement around data itself is the larger driver. From the podcast (emphasis mine):
“I think the vast majority of pre-training data improvements are from science on better understanding what data sets are good and schleppy labor on figuring out how to filter down.
So my view is that improvements of the form of, like, OpenWebText to FineWeb, that improvement is better described as an algorithmic improvement of the sort that you can study with some GPUs, and you don’t need human expert data to do that. Now, there’s a different effect which we could talk about, which is that maybe the internet in 2026 is more of a fertile ground for training data than the internet in 2018. There’s also been an effect where there are just more humans posting on the internet, so there’s more data to harvest. My sense is that that effect is going to be quite a bit smaller than the effect of humans knowing better how to curate the data, having better scrapes, knowing how to process those scrapes better — this sort of thing.”
Greenblatt’s arguments certainly gave me a pause because my prior was a bit different and likely much closer to Dwarkesh who also appears to think human expert data played a critical role in the model’s recent trajectory.
Bi probably agrees with Greenblatt since he decomposed where human knowledge comes from: a loop of proposing tasks, learning existing knowledge, thinking, getting feedback from the environment, and distilling the findings back into knowledge and wondered aloud in his talk which steps AI can accelerate. His answer is basically nearly all of them, including proposing the tasks in the first place. Greenblatt essentially makes the same claim retrospectively: RL environments improved over the last two years mostly because labs learned what to build and used enormous amounts of AI labor to build it.


Another interesting observation by Greenblatt was that machine learning (ML) is a “shallow” domain compared to math and given that even in math we are transitioning “from an era of proof scarcity to an era of proof abundance”, automating much of ML may prove to be lot more amenable. From Greenblatt:
I think ML is a very shallow domain relative to math. In math, there was much more of a thing where you find some true deep abstraction, and if you really understand that thing, which is hard to understand, then you get somewhere. Whereas I feel like the things that are the equivalent of that in ML are really dumb bullshit. Like with scaling laws, come on guys, we can explain scaling laws really quickly. I think the deepest and most important concepts in math, for example, don’t have the property that you can really understand the underlying thing and why it matters in a very short period of time.
My sense is that some domains are structurally different in terms of how they operate and how much they depend on deep abstractions. Physics and math are much more on the side of being very far on the deep, hard-to-come-up-with-ideas side, whereas I think ML and most other domains are much more amenable to hill climbing. That’s my sense of how this will go in the future.
Even in cases where there has been some breakthrough in AI, oftentimes in retrospect it looks like a big bottleneck to making that breakthrough happen was getting all of the micro details and mungy intuition right. An example of this is training AIs to be good at reasoning and chain of thought, doing RL on chain of thought. It looks like you probably could have done RL and chain of thought on GPT-3 and gotten kind of interesting results on math if you had really scaled it up and done a good job.
Bi also made the point that learning from environment interaction is efficient wherever a perfect simulator exists (coding, math etc.) and fundamentally blocked where simulation is impossible or the sim-to-real gap is simply too large ( for example biology, and experimental physics). Given that context, automating AI research doesn’t seem nearly as outlandish.

Gavin Baker in a recent reply to an Anthropic researcher mentioned, “I do think the computational efficiency of humans (I’m impressed that your brain runs on only 15-20 watts tbh) means that humans will be economically useful for the foreseeable future *even* in a fast takeoff, AGI maximalist scenario.”
Indeed, the efficiency of humans was also highlighted by Bi during his presentation (see a bunch of related slides below).



However, the amazing efficiency of homo sapiens is also a stark reminder that nature has already shipped a general intelligence that runs on less power than a dim lightbulb. Today’s models need a building full of GPUs and most of the written internet to often do less. Does that indicate something “special” about us or is it a measure of how inefficient our current approach still is? Bi’s bet is that the biggest waste sits in how models learn. If someone fixes that, the cost of a given level of intelligence can fall by orders of magnitude. Of course, I am not in a position to know or predict whether this is at all fixable or even if it is, when that may happen.
A peer recently praised me to help him improve his understanding of the AI landscape through my work at MBI Deep Dives. I jokingly mentioned to him I’m glad that you feel that way, but ironically the more I study AI landscape, the more certain I become that I need to hold every opinion related to AI very loosely. Such a frame of mind doesn’t inspire a lot of confidence in my own mind that I can see too far ahead. Investors are trying to price AI landscape based on near-term trajectory of respective companies, but given how fast things can alter in the AI landscape, it is hard not to feel that betting for or against this trade carries a monumental risk either way.
Subscribers get the daily journal and five+ years of Deep Dives, i.e. full-length analyses with financial models on 65+ companies. The daily is just how I think out loud between the Deep Dives!
Current Portfolio:
Please note that these are NOT my recommendation to buy/sell these securities, but just disclosure from my end so that you can assess potential biases that I may have because of my own personal portfolio holdings. Always consider my write-up my personal investing journal and never forget my objectives, risk tolerance, and constraints may have no resemblance to yours.
My current portfolio is disclosed below: