Wealth Systems

Wealth Systems

The Great Repricing of Intelligence

What open weights, local hardware, and collapsing inference costs mean for investors and entrepreneurs

Matt McDonagh's avatar
Matt McDonagh
Jul 26, 2026
∙ Paid

My life orbits a single obsession: maximizing return on time.

That focus originally led me to investment banking, and ultimately drove me to frontier AI and agentic engineering.

Way back in the good old days (2010) while building a hedge fund focused on automating complex financial analysis using machine learning, I realized that while AI acts as the physics of modern value creation, data is entirely responsible for its logistics.

Scale demands more than just a well-built corporate brain, you have to master the nervous system.

Now, operating as an Agentic and Data Engineer, I build that critical infrastructure. I design AI agent operating systems that fuse cutting-edge models with rigorous business strategy. From Fortune 500s and legacy professional services firms to family offices and rapid-growth startups, I engineer the solutions that actually make AI operational in the real world.
I provide advisory services to private equity firms, investment banks, and enterprise B2B firms on AI market positioning, demand trends, AI Governance, AI Evaluation and AI Implementation strategy.

This hands-on building is my greatest asset as a technology investor. Because I spend my time engineering the foundation, I don't just evaluate the surface-level of a new AI company. I audit the data logistics, the agentic architecture, the revenue systems maturity and the viability of the product.

Building the future gives me the lens I need to invest in it.

Let me tell you what I am most excited about: the cost curve of intelligence.

It’s dropping even faster than the most bullish techno-optimists ever dreamed… and the intelligence explosion is accelerating.

The Great Repricing

The most important economic fact about AI is not that the models keep getting smarter.

It is that useful intelligence is getting cheaper faster than markets know how to price.

For the last several years, the AI story has been told through scarcity. Frontier models required rare talent, giant clusters, enormous power, and billions of dollars. The companies that could assemble those inputs appeared to own the future. Intelligence looked like a scarce industrial product that the rest of us would rent through an API.

That story is still true at the very edge of the frontier.

It is becoming less true everywhere else.

Open and open-weight models are improving. Chinese labs are forcing the price-performance curve lower. Quantization is shrinking models. Inference engines are getting faster. Apple, AMD, NVIDIA, and others are putting more memory bandwidth and AI-specific compute into personal machines. The software around the model, called harnesses, is learning how to route tasks, hold context, call tools, test results, and escalate hard work.

This changes the ownership structure of intelligence.

The frontier remains centralized.

The useful frontier is moving outward.

That split will reshape venture capital, public markets, company formation, software pricing, labor, and the balance of power between platforms and users. The opportunity is not simply to invest in more AI.

We need to think about what becomes scarce when intelligence does not.

Two Curves Are Moving at Once

AI is developing along two curves that look contradictory.

The first curve is the industrialization of frontier training. Global AI compute capacity has been growing at extraordinary speed. The 2026 Stanford AI Index estimates that global capacity grew 3.3 times per year from 2022, reaching 17.1 million H100-equivalents. NVIDIA accounted for more than 60% of that capacity, while the most advanced chips remained heavily dependent on one foundry in Taiwan.

This is concentration economics. More capital. More power. More networking. More advanced packaging. More geopolitical risk. The next frontier model may cost more to train, not less.

The second curve is the commoditization of inference.

Stanford found that the cost of using a model with roughly GPT-3.5-level benchmark performance fell from $20 per million tokens in late 2022 to seven cents by late 2024, a decline of more than 280 times. It also found that the smallest model to cross a particular benchmark threshold shrank 142-fold in two years.

Think about the implications of that: the dollar cost fell. The model got smaller.

The hardware got better. And the ecosystem learned how to squeeze more work from the same weights.

This second curve is now landing on the desk. AMD is shipping laptop-class systems with 128 gigabytes of unified memory and demonstrates meaningful local inference without a cloud dependency. Apple’s M5 Max supports up to 128 gigabytes of unified memory and 614 gigabytes per second of memory bandwidth. Apple now tells developers to run models entirely on-device, with no server dependency or token bill. NVIDIA positions 128-gigabyte personal systems as capable of running models with up to 200 billion parameters.

This is the promise of generating hundreds of tokens per second on your local machine without any need for API bills, internet access or the risk of sending your intelligence to the frontier.

Marketing claims are not the same as production economics. A model fitting in memory does not mean it runs fast enough, reliably enough, or cheaply enough for every workflow. But the direction is obvious.

What required a data center moves to a workstation.

What required a workstation moves to a laptop.

What required a premium model moves to a smaller one.

This is how computing diffuses.

Intelligence Is Becoming a Capital Good

In When Frontier Intelligence Comes Home, I argued that local inference turns intelligence from a metered service into installed capacity. That distinction has larger economic consequences than it first appears.

An API is rented intelligence. You pay each time the system thinks.

A local model is a capital good. You buy the machine, load the capability, and decide how hard to run it.

The cost does not disappear. Hardware depreciates. Electricity matters. Engineers must maintain the stack. Models become obsolete. Local systems may be slower than cloud clusters and less capable than the best closed models.

But the incentive changes.

With an API, every additional attempt creates a variable cost. With owned compute, idle capacity is the waste. Once the machine exists, the rational move is to find more work for it.

Run another review. Test another design. Reconcile the archive again. Give every support ticket a first-pass investigation. Let three models attack the same problem and make a fourth compare the answers.

This is the same logic that turned cheap bandwidth into streaming, cheap storage into photo archives, and cheap compute into software everywhere. Falling unit costs rarely reduce total demand. They create new behavior.

The price of a unit of cognition falls.

The quantity of cognition demanded explodes.

That is why the simple “AI will lower costs” frame is too small. AI will lower the cost of existing tasks, but it will also make millions of previously uneconomic tasks worth doing. The result will be deflation inside the task and expansion across the system.

Cheaper intelligence will not only automate the work we have.

It will finance the work we have never been able to afford.

The Seed Investor

The Moat Moved

Seed investing is my favorite form of investing. For a seed investor, the first consequence of the forces discussed above is brutal: access to a strong model is not a moat.

It may not even be an advantage for long.

A startup can build on a frontier API today, shift routine traffic to an open model tomorrow, self-host a specialized model next quarter, and replace the entire model layer a year later.

OpenAI-compatible interfaces and model routers make substitution easier. The open ecosystem is becoming too large to dismiss; Stanford counts 5.6 million AI-related projects on GitHub, while uploads to Hugging Face have tripled since 2023.

Not all of that is truly open source. The Open Source Initiative’s definition requires the freedoms to use, study, modify, and share, along with access to the relevant code, parameters, and data information. Many important releases are better described as open-weight: downloadable and adaptable, but governed by licenses and disclosures that fall short of full openness.

The legal distinction matters.

The economic direction matters more.

Model capability is becoming substitutable. The seed investor must underwrite what survives the substitution.

That means proprietary workflow data. The right to act inside a customer’s system. Evaluations built from real completed work. Distribution into a valuable niche. Trust earned through reliability. Feedback loops that improve the product each time work passes through it.

The moat is no longer intelligence access.

The moat is intelligence allocation.

This also changes venture math. AI makes it cheaper to build software, conduct research, produce content, support customers, and operate the back office. A ten-person company can attempt what once required fifty. Seed rounds can last longer. Founders can reach revenue with less capital. Some companies will become extraordinarily profitable before they become large employers.

That sounds unambiguously good for venture investors. It is not.

The same tools that make one portfolio company more efficient make it easier for twenty competitors to appear. Product cycles compress. Features copy quickly. Technical differentiation decays. Revenue can grow faster while durability gets weaker.

Seed investing therefore becomes less about asking, “Can this team build the product?” More teams can.

The sharper questions are: Can they become the system of record for a workflow? Can they capture unique outcome data? Can they earn permission to take consequential action? Can they distribute into a market faster than the product is copied? Can they compound while the underlying model is replaced five times?

The best seed companies will not wrap a model.

They will own a control point.

The Public-Market Investor

Follow the Bottleneck

Public markets face a different problem. A tricky one.

They are trying to value a transition in which demand can surge while unit prices collapse.

The instinct is to choose a side: centralized or local, cloud or edge, closed or open. That is the wrong frame. Both curves can win at the same time. Here’s how.

User's avatar

Continue reading this post for free, courtesy of Matt McDonagh.

Or purchase a paid subscription.
© 2026 Matt McDonagh · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture