Rendered at 06:10:36 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
tristanj 8 hours ago [-]
The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations. Anthropic is valued at $1.2T and OpenAI is targeting $850B. These astronomical valuations were built on the premise that these labs would generate massive profits from premium API pricing, but the Chinese labs are completely undercutting this strategy by releasing excellent open models for free. If the frontier labs are forced to cut prices and join the race to the bottom in token prices, these valuations are unjustified, and VCs will face enormous (paper) losses.
mediaman 7 hours ago [-]
The (quite excellent) article discusses several of your points. If you haven't read it, I recommend it.
- Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.
- Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
- The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.
- This is because US labs are leading on cost efficacy of inference ($/task)
- Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
- With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.
larrysalibra 3 hours ago [-]
Ben's article "distills" down to 2 reasons that US frontier labs shouldn't be "afraid":
1. US frontier lab unit economics are better
2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.
For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.
For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.
And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.
And every improvement in model capability makes it increasingly easier to make your own tools.
He’s glossing over the reason they are not: 90% profit margin of Nvidia. Power is only a small part, single digit, it will eventually matter but does not really today.
What is the cost of AI? The single largest ingredient is Nvidia profit margin.
Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.
Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.
Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)
aurareturn 2 hours ago [-]
Cost of electricity isn’t a long term advantage in my opinion. Private companies will figure it out.
What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.
I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.
Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.
wbadart 35 minutes ago [-]
Seems like most popular harnesses, including codex and Claude code, support Agent Skills (an open spec for skill formatting/ organization): https://agentskills.io/clients
Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)
culi 1 hours ago [-]
> Private companies will figure it out.
Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy
golem14 2 hours ago [-]
I'd really love to see the evidence on this!
tvink 22 minutes ago [-]
I think calling opencode the trend is naive. This not what is being run on company time.
hack1312 10 minutes ago [-]
OpenCode is absolutely used on company time.
Computer0 1 hours ago [-]
In a corporate setting yes Opencode all the way. However in a non corporate setting I am getting $3000 of api usage a month for $100 at Anthropic and only use open code for the smallest cheapest tasks
pishpash 2 hours ago [-]
More basically, production cost matters only if inference is priced at commodity prices. That's not what VC's signed up for, which is rent-seeking.
ehnto 3 hours ago [-]
US running costs are higher than in China, because the US lags behind in energy, has higher real estate costs, and wage costs are higher.
Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)
Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.
The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.
Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.
All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.
Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.
analyte123 3 hours ago [-]
The US does not lag behind in energy. Industrial electricity prices in most places in the US are competitive with China, or even cheaper.
ehnto 28 minutes ago [-]
Apologies, I meant ability to deploy new generation. You're right that areas of the US are cost competitive.
This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.
zx8080 3 hours ago [-]
Links please.
analyte123 1 hours ago [-]
IEA’s chart for “final electricity price for large industrial customers in energy-intensive industries” [1] shows the US cheaper than China, but this is using data from Texas as representative. I’m not sure how they back this representation up, or how different this is from GPP “business” prices. So “most places in the US cheaper than China” may be wrong, but certain states or regions are at least competitive.
This seems to show that even with China taxing industrial electricity to subsidize household electric bills it's still 30% cheaper for industrial electricity in China. [0] Those same taxes make household electricity about half the price of US power.
If I had to wager why, I'd say it's due to embracing solar on massive scales recently. Only a few years ago the US was competitively priced.
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]
I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.
It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?
I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.
gruez 6 hours ago [-]
>What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively
Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.
If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.
striking 6 hours ago [-]
I'm arguing we can't trust retail prices because the marginal pricing isn't meaningfully connected to it anyway.
But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.
And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
striking 5 hours ago [-]
The fixed consumption plans are offering several times their worth compared to API pricing with completely free cache reads: https://she-llac.com/claude-limits
I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.
> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?
Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.
c0brac0bra 4 hours ago [-]
The 33% discounted competitors have no non-retention policy
ycui7 3 hours ago [-]
discounted competitor could cheat. they can offer subpar model response and sell it as deepseek-v4. it is uneconomical to prove inference providers are cheating, so they get away with it. cheating inference provider does not care if their customers stay.
appplication 6 hours ago [-]
I think there are some really interesting thought there, but I’d challenge some of this:
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).
The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.
Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.
> This is because US labs are leading on cost efficacy of inference ($/task)
It’s possible, but I would need to see better data on this.
>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.
senderista 4 hours ago [-]
Amazon is not "meaningfully in inference"? Bedrock seems to have a ton of enterprise customers, some of which would never trust the AI labs themselves with their data but will trust Amazon.
appplication 4 hours ago [-]
Relative to their other compute or other players inference, not as significantly. Though yes, they certainly have some share.
resonious 21 minutes ago [-]
> for frontier work.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
SoftTalker 3 hours ago [-]
> Models are not free. Downloading them is free. Running them is not.
Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.
aurareturn 2 hours ago [-]
I think the point here is that it takes the same hardware to inference an open source model as OpenAI/Anthropic inference their models.
IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.
So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.
pishpash 2 hours ago [-]
No they don't. Are OpenAI and Anthropic charging at cost, or wish to? No.
aurareturn 2 hours ago [-]
If you have more efficient models, you can reduce price and grab market share or keep price and increase your margins.
Over time, this compounds. More profits means more investments. More market share means more control.
lemax 7 hours ago [-]
But this assumes Chinese models will not achieve token cost optimization. Intelligence needs are fairly flat for many tasks, and the Chinese models have caught up on this front. Next they achieve greater token cost efficiency and we don’t need OpenAI.
overfeed 4 hours ago [-]
That the author doesn't acknowledge the relentless R&D efforts DeepSeek has been plowing into optimization, and giving a default win to OpenAI/Anthropic on the supposition that they've been serving models for longer is a black mark against the article.
I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.
VulgarExigency 6 hours ago [-]
The model that is most optimized around token cost is, in fact, Chinese. DeepSeek is astoundingly cheap by default, but if you use it from Reasonix (the harness optimized around its cache), it becomes even cheaper.
ryeguy 2 hours ago [-]
I keep seeing mention of the cache, what's special about it? All frontier llms have prefix caching, what is special about deepseek's approach?
ChaitanyaSai 3 hours ago [-]
The thing I do not understand here because it seems obvious: AI will be a commodity market and you simply cannot have a large PE multiple. So the valuations imagine a global commodity monopoly or duopoly coupled with the increased intelligence still disallowing other suppliers from becoming competitive? Without any network effects to help?
horacemorace 3 hours ago [-]
Perhaps to moneymen the difference between “ChatGPT” and the technology behind it isnt’t obvious. I’ve been very surprised at how few otherwise smart people are completely in the dark about how capable current models are.
As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.
didibus 2 hours ago [-]
China is working on the whole supply chain though and they're willing to compete on razor thin margins. Just look at EVs. They build great cars but the competition is so aggressive that investing in any one Chinese EV company isn't exactly an amazing ROI.
I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.
AnthonyMouse 3 hours ago [-]
> Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.
This is the story for Nvidia/AMD or cloud providers rather than OpenAI.
> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
It seems like there would be problems with this on both ends.
For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.
Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.
And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.
Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.
marcus_holmes 3 hours ago [-]
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I notice that the article, and this discussion, hasn't mentioned or considered local models.
We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.
If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?
nunez 9 minutes ago [-]
This will matter A LOT more after Apple gets serious about integrating AI into the OS.
gmerc 1 hours ago [-]
Yea but at those rates VCs will never make their money back. Because Deepseek and friends keep releasing the inference optimisations to everyone instead of holding them back to pay their investors.
2 hours ago [-]
abernard1 6 hours ago [-]
" - The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6."
This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.
But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.
Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.
mediaman 4 hours ago [-]
You're correct, but that's a different market segment and not the market GLM 5.2 and its peers compete in.
The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.
From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.
But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.
That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)
ignoramous 5 hours ago [-]
> "The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6." This is a flatly false statement.
It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".
Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.
[0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).
abernard1 5 hours ago [-]
Fair. Too strong a statement.
But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.
It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.
I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.
Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.
tristanj 6 hours ago [-]
I did read the article, but it misses the core issue entirely, and it's why I shared my comment to begin with. Look at the cost-per-task benchmarks from Artificial Analysis https://artificialanalysis.ai/models?cost=cost-per-task
Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.
But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.
* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.
* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.
* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].
OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.
[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware"https://x.com/rohanpaul_ai/status/2079027313455550839
isodev 4 hours ago [-]
Imagine how cool it would be if actual competition prevents Anthropic or OpenAI from becoming an Apple/Google kind of cartel. I don’t care if it comes from China or not.
AlexCoventry 3 hours ago [-]
You don't care if you're running models created by people subject to a Stalinist dictatorship, which could have forced them to train the models to report on you or otherwise run amok if they encounter an arbitrary trigger text?
I would say the weirdest thing about the article was that it was entirely in terms of business value. But commercial businesses are at best just a stop-gap measure, to a Marxist.
culi 1 hours ago [-]
And here I was believing the Red Scare was for the history books
rootsudo 14 minutes ago [-]
Same, if only they would take a look at where their iPhone and MacBook was manufactured. Oh wait, you have a thinkpad? Isn’t that Lenovo?
__MatrixMan__ 2 hours ago [-]
Are you saying that the US companies are above that kind of behavior?
platevoltage 15 minutes ago [-]
Are the Marxists in the room with you now? Do you need assistance?
JSR_FDED 2 hours ago [-]
Oh give me a break.
Using a Chinese LLM will not put a Marxist under your bed.
Did you know Gemini is shockingly bad at French poetry? Hasn’t stopped me for using it for all other tasks though.
bornfreddy 1 minutes ago [-]
> Using a Chinese LLM will not put a Marxist under your bed.
Interesting tangent - it might be taught to introduce stealthy backdoors in your company though. Maybe even across multiple PRs where each session puts a small chink in the armor, and together they allow unlimited access to the attacker who knows about them.
After all, LLMs are mostly black boxes. How comfortable would you be running a Chinese compiler?
killingtime74 2 hours ago [-]
Good old FUD
testbjjl 3 hours ago [-]
We both know the answer. Write offs. If your fund was not in AI heavy you’d have no investors.
wordpad 6 hours ago [-]
I think everyone understands models will be a commodity.
Its the user base (with ads and upselling) and proprietary wrappers which will make money for typical customer.
Even enterprise customers arent going to be spending a lot on tokens. Once labs no longer have to subsidize trainings tokens costs will drop 10x and once models get burned on chips costs will drop 10x more and you physically won't be able to burn significant number of tokens unless you're deliberately trying to.
skeptic_ai 5 hours ago [-]
[dead]
walrus01 1 hours ago [-]
I would say it's not just the VCs but the various other entities that will be left holding the bag of debt if the AI-fueled datacenter construction boom/bubble pops. For a list of large and well known projects and their scales:
I had plenty of gains by holding investments which Goldman Suchs were making doom statements about.
rootsudo 12 minutes ago [-]
Always buy opposite, buy low and sell high.
Too many people here buy high and sell low.
When Goldman makes statement, half the time it is prepped by an associate or two that has minimal experience and goes through a MD that enjoys the gloom and doom. That’s why they publish. Goldman makes money on both sides of a trade.
rvz 5 hours ago [-]
> The people who are most afraid of Chinese models are the VCs who poured into Anthropic and OpenAI at astronomically high valuations.
Correct. These chinese labs has proven that having just the model is not a moat, and the safety concerns were all just attempts at regulatory capture.
This is why labs like OpenAI and Anthropic are panicking and are racing to the exit before their valuations start being questioned.
eeiei 4 hours ago [-]
[dead]
onesociety2022 6 hours ago [-]
But there a ton of other VCs who poured money into SaaS businesses. They have the opposite incentive. They want tokens to be cheap like a commodity so the value accrues in the SaaS/app layer.
hamandcheese 5 hours ago [-]
Cheap tokens only benefits SaaS that depends on AI. Otherwise, cheap tokens means it is only more cost effective than it already is to cut out the SaaS and build instead of buy.
Ericson2314 3 hours ago [-]
Yeah maybe *both" sass and training companies are wiped, imagine that!
pishpash 2 hours ago [-]
Maybe AI is deflationary, in the right hands.
fnord77 5 hours ago [-]
I guess I shouldn't try to buy shares of OpenAI on the private market...
ProAm 2 hours ago [-]
tbh they are floundering even to regular investors. They are trying to give the US gov 5% of the company so they become 'too big to fail' but they are in trouble.
airstrike 8 hours ago [-]
they will likely suffer enormous real losses too, not just paper, though not as enormous
for VCs, breaking even is losing
cyanydeez 6 hours ago [-]
mmm, the chinese models are also working on local GPUs at consumer grades. so theyre not just drainig cloud moats.
vrm 6 hours ago [-]
good luck running a 2.4T model on any local hardware. it’s not gonna happen. the arrow is to specialized hardware at least for the smartest models
matheusmoreira 5 hours ago [-]
I have hope it'll happen one day, even if not now.
nekusar 5 hours ago [-]
Already is possible. On a machine with 32GB ram, and NO gpu. Just need a large SSD or NVME. Streams from disk to memory.
Well, then let's hope you're wrong and the AI bubble won't also blow up private equity and wipe out people's retirement funds...
lorecore 7 hours ago [-]
Good. Over the past few years, VCs have proven that they’re warmongering psychopaths. Hopefully China puts every last one of the Palantir/Flock/Anduril class out of business.
happypappy123 6 hours ago [-]
Lol, chinese surveillance makes flock blush
lorecore 6 hours ago [-]
I’m 100% certain that China won’t be sending any goons to my front door.
lmz 4 hours ago [-]
Not that I'm anti-China, but their companies would have no qualms selling surveillance tech to your local gov too if they could. I'm sure Palantir, Flock, and other US companies would sell to China too if they could.
yonaguska 4 hours ago [-]
This is true, but there is another foreign country that can send people to your door. What's to stop China from eventually buying that type of influence over our govt officials?
platevoltage 12 minutes ago [-]
Given the nations that have actually bought influence over our government officials, China would be an upgrade. I doubt the president needs another decked out 747 though.
Revanche1367 4 hours ago [-]
If our govt officials are willing to betray their citizens for money to China, what’s the point of preferring to give them power over us instead of giving it to China?
phendrenad2 4 hours ago [-]
VCs are just pass-through investors, the money comes from billionaires. And when billionaires face losing money, the whole system re-arranges itself to stop that from happening.
reinitctxoffset 2 hours ago [-]
[dead]
alex1138 3 hours ago [-]
I'm a civilian, not a VC. In my own case, I'm worried how many things pass through the CCP. How censorship of mentions of Tiananmen Square is something they're quite interested in
riskd 2 hours ago [-]
How exhausting.
alex1138 2 hours ago [-]
Why? Why is this not a concern?
JSR_FDED 2 hours ago [-]
Because you’re not that important. I don’t mean that in a mean way, just that if you’re a serious player in national security or something like that, you’re already not using Gmail, let alone ChatGPT.
platevoltage 12 minutes ago [-]
Claude refuses to call Trump a Fascist.
wxw 8 hours ago [-]
> It’s striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users.
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
Aurornis 6 hours ago [-]
For personal use I agree.
For companies, these decisions are very sticky. Companies go through a lot of red tape to get anything purchased and approved, then they discourage change because it's a lot of work.
So the product that gets a foothold in a company sticks for a long time.
Then a couple years later a sales person convinces an exec that they can save some money by switching, so the switching game begins. Not necessarily motivated by the better product, mostly the price. My wife's company keeps switching their tools out from under everyone every year or two. Just when they get everything stabilized and everyone familiar with the new tool, some new contract is signed that moves them all to some other company's suite.
exhilaration 3 hours ago [-]
I'm confused, I work at a big giant Fortune 500, we all get GitHub Copilot subscriptionsn
- we can switch between OpenAI and Anthropic models with just a click in Visual Studio. There's no stickiness at all. They just made us go through a training after the price hikes about how to choose between models for the best cost/benefit ratio.
sothatsit 2 hours ago [-]
The models are not what is being discussed here, it is the harnesses. That is, Claude Code, Codex, and what you use, GitHub Copilot. I suspect there would have to be strong reasons for your Fortune 500 company to switch away from Copilot.
Similarly, I have made no ground in arguing to try to get Codex at the company I work for, which got Claude Code a year ago and sees no reason to go through the whole process of setting up any alternatives when Claude Code already works and is at the frontier.
Aurornis 2 hours ago [-]
GitHub Copilot is the sticky product in your org.
You can choose a selection of different models within it, but you're not using Codex or Claude Code.
thaeli 2 hours ago [-]
Same here, except we just defaulted everyone to Auto and expect the percentage of tokens spent via the auto router to be high.
rohansood15 5 hours ago [-]
Companies have learned their lessons on stickiness with cloud providers. Every enterprise has a multi-provider strategy now.
andersonpico 6 hours ago [-]
Every company that I've worked with that provided models internally did so through LiteLLM and offered both Anthropic and OpenAI models so it was trivial to switch between them.
blfr 6 hours ago [-]
Most companies just get you a Claude team sub and maybe a couple of skills.
linkregister 3 hours ago [-]
The parent poster is almost certainly talking about inference within workflows and not for interactive coding agents.
stingraycharles 6 hours ago [-]
We only get Copilot. I’m not very happy.
AgentME 4 hours ago [-]
What do you find worse about it? I've been switching between it, Codex, and Claude Code to try to compare them, and my only conclusion so far has been that it's nice that Copilot has both OpenAI and Anthropic models as options.
Aurornis 2 hours ago [-]
I like all the different comments in this thread saying that most companies do X, where X is a different answer from each person: LiteLLM, GitHub Copilot, Claude Code.
HaloZero 3 hours ago [-]
I imagine the play here is going be connectors. Can you get slack to avoid integrating with anyone other American ai providers, same with Google suite, etc etc.
linkregister 3 hours ago [-]
It's almost trivial to create a custom Slack application wrapping your desired harness running in a container on your organization's k8s cluster. Likewise with MCP support. These are already open.
Aperocky 43 minutes ago [-]
I use a mixture of claude code and codex and kiro as my swarm.
They communicate through my own harness, and it's working pretty well so far. claude code is being overtaken by codex however because I noticed lately the accuracy of the latter is the best.
happypappy123 8 hours ago [-]
Their form and function have basically converged, sometimes I will open one up and confuse it with another
SOLAR_FIELDS 7 hours ago [-]
Which would imply that these things are fast becoming… checks notes… a commodity?
bushbaba 5 hours ago [-]
agreed, my F500 company switched off claude code to copilot in 30 days. All 5k+ engineers. That is the fastest migration i've ever witnessed. This includes switching all our agents from Claude SDK to Copilot SDK.
philstephenson 2 hours ago [-]
As a Hacker News user and commenter, you are not the type of user he’s referring to.
nl 6 hours ago [-]
Have you ever worked with a non-programmer and helped them setup their AI workflows?
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
favouritemartin 6 hours ago [-]
Skills are quite interoperable, and you can easily ask Codex / Claude to help you with switching the MCP connectors or any other things specific to your previous workflow. It's been quite low friction in my experience.
nl 4 hours ago [-]
I know someone who runs AI training.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
andrewf 6 hours ago [-]
It strikes me as like setting up an IDE. People have preferences, switching is possible, but there are advantages to saying "we are a Visual Studio + Resharper shop" or "everyone uses IntelliJ to work on this project".
trollbridge 5 hours ago [-]
Yes. I taught the non-programmer to ask the harness to set up things like MCP connectors.
cyanydeez 5 hours ago [-]
we have AI. WHAT is it good for if a harness cant just take a api endpoint and some permissions and duplicate.
its so distracting seeing these types of confision.
every plugin is already just multimodaling their targets.
IAmGraydon 7 hours ago [-]
Same here. I flip flop between them. Most people I know who have access to both, technical or not, are doing the same. They’re just too close and sometimes one does what you want better than the other.
solumunus 8 hours ago [-]
I think they stickiness is less about the difficulty of switching and more about the lack of desire. I’ve been using Claude since day one, it works well and I’m happy, I like it. I’m sure Codex is good too. Switching from one to the other certainly isn’t going to be a game changer, the discourse shows me the differences are marginal.
Probably the only reasons I would seek change are economical.
jjfoooo4 3 hours ago [-]
A sticky product is one that switching away from creates a major hassle. Which means the user will pay more to avoid said hassle.
“I don’t really have a strong preference between the two” is another way of saying “the product isn’t sticky”, which is another way of saying “this provider has very little room to increase margins”
solumunus 24 minutes ago [-]
No, I don’t think so, and searching seems to confirm my view. Inconvenience of switching is just one aspect.
There’s little difference between Coke and Pepsi and the barrier to switching is nil, yet clearly the products have stickiness. People have slight preferences and become familiar with the brand and then engagement becomes habitual.
The effects on margins are irrelevant to this.
mediaman 7 hours ago [-]
Convergence in coding makes them highly substitutable. But I could see harnesses configured for different purposes -- let's say, a harness for creating teaching plans -- being able to cater to its audience better than a coding harness. Maybe it's got tools to plug into standardized curricula, what the lesson books will be, what other lesson plans the district's teachers have made, etc., which could be done in a clunky way in a regular harness but could be streamlined.
sergiotapia 5 hours ago [-]
My same progression here. I started with ChatGPT website, then Anthropic website, then Cursor, then Windsurf!, then claude, then opencode, then ohmypi, then codex, finally back on Cursor now because I think they cracked the UX for what great dev looks like. The grok 4.5 fast model + cursor ergonomics is insanely good!
The cost of me moving around these different AI models and harnesses was pretty much 0.
kinj28 3 hours ago [-]
I am afraid — if Chinese models go mainstream it has a clear way of pushing its narrative way beyond its otherwise borders. More like a Trojan horse it is for the Chinese.
Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question
Is this something that is more true of a Chinese model than any other model of a different national origin?
Genuine question: generalized up from individual models to “models from country X”, is there any country that doesn’t have this exact risk?
itake 2 hours ago [-]
In the USA, multiple political parties balance each out other.
In China, there is 1 party. 1 view. 1 definition of the Truth.
TripolitianFish 40 minutes ago [-]
I’d love to live in the USA you’re talking about friend.
This is just oriental despotism paranoia, whatever cutsie repetition slogan you come up with is not a serious argument.
kinj28 39 minutes ago [-]
I would think everything boils down to source of funding and their narrative should get pushed!
jihoons 2 hours ago [-]
But having a cheaper, open-source model provides alternatives to the market for proliferation and adoption of the tech in scale
kubb 2 hours ago [-]
What if the Chinese model has a point?
xlmnxp 2 hours ago [-]
Same thing happen for western models, try to ask about Gaza genocide and see for what side it will stand
itake 2 hours ago [-]
What am I trying to understand about the western models on the term "Gaza genocide"?
I asked Grok "Tell me about the gaza genocide" and it write a IMHO balanced answer comparing why genocide is and isn't the right term. [0]
ChatGPT 5.6 Sol only explained why people call it a genocide and did not go in as in depth as Grok did for why people don't agree with the term. [1]
The only unsaid response (to me) here is the model should have declared that it was not a genocide, and because these models explain why it was a genocide, they are bad?
So all the models you tried disagreed with the United Nations ruling?
In international law, that is the ultimate authority of what is/isn't a genocide. It's objectively, legally speaking, a genocide.
itake 34 minutes ago [-]
I have a couple questions.
1/ I don't see in the responses where the model says it is or isn't a genocide. Can you share the snippet from each, I included the logs above?
2/ I can't find a source on the UN ruling that you mentioned. I am not interested in the findings of an investigative body, just the official UN ruling. Can you share? ChatGPT (and myself) can only find this [0], which is a second round of written submissions.
This kind of response is predictable, since these models are tuned to align with one side of the issue. A recent example: Grok answered a similar question and was suspended shortly afterward [1]. Furthermore, a UN commission has formally concluded that it constitutes genocide [2] — a fact that Western AI models rarely mention, let alone link to directly.
I operate an analytics site (pretty big one B2B where client's backend feeds data into our system), and we see tons of traffic originating from northwestern China (Xinjiang) from Shenzhen Tencent Computer Systems Company Limited.
There are also half a dozen other companies from China continuously hammering our clients’ websites.
I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.
Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region
credit:
'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA
HN user: lopoc
Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.
Assuming that China only distills is a huge mistake.
It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.
Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.
The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.
Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.
This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....
There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
ballon_monkey 3 hours ago [-]
The 2 things people need to remember:
1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
shunia_huang 1 minutes ago [-]
But Dario said (and maybe more ppl) that Chinese models are just distillation of their model and training data, so I guess your first point is invalid?
makeitdouble 2 hours ago [-]
On (1), models are an aggregation of large volumes of data sources, whatever the culture producing them, you'll get the average bias of that culture.
We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.
MintPaw 10 minutes ago [-]
I see your point, but this isn't exactly true, it only takes a single person to bias a model by deleting specific training data or over training on certain facts.
riskd 2 hours ago [-]
It’s perfectly fine for the “West” to influence the world though, right? Or is it only a problem because… they’re Chinese?
hetman 57 minutes ago [-]
Why would it be a problem that they're Chinese? It's a problem because their country is ruled by an authoritarian regime. The West, by contrast, doesn't have a singular arbiter of truth, it has a plurality of perspectives (as much as certain interest groups would really rather that was not the case). It's not perfect, but those of us who have memories of living under authoritarian regimes can tell you there's no comparison.
youre-wrong3 1 hours ago [-]
Have you got examples of GPT/Claude/Grok influencing people?
hetman 44 minutes ago [-]
I think there's no denying they each have an ideological bent (as is their right as private companies). I have had Copilot deny me access to historical information on ethical grounds, even though I don't think anyone would find it controversial (clearly it was overtuned, GPT and Claude had no problem answering the same question). What is different though is that they are each allowed to have their own perspective instead of a singular mandated one.
pwn0 2 hours ago [-]
1. Get the free Chinese model.
2. Jailbreak it
3. ???
4. Profit?
bigyabai 3 hours ago [-]
1) I'm not using AI to bicker over fringe political shibboleths.
2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.
sanex 3 hours ago [-]
Different county different feds both spying I'm sure.
margalabargala 2 hours ago [-]
Sure but if you are not Chinese or in China, then the chances of negative consequences to you from the Chinese feds is vanishingly small due to lack of ability to do anything that affects you.
Meanwhile if you are in the US, DHS has already subpoenaed social media sites looking for people who made anti-ICE posts and I can't imagine they consider subpoenaing AI conversations off limits https://www.nytimes.com/2026/02/13/technology/dhs-anti-ice-s...
walrus01 3 hours ago [-]
> 1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
One of the interesting things is that through a fairly rudimentary process which is being done by 3rd party amateurs who've downloaded the open models, models like Qwen 3.6 35B-A3B (or 27B) can be fully 'uncensored' when turned into GGUF files.
I have an uncensored Q8 version of Qwen 3.6 35B-A3B here that will very happily output information about Tiananmen Square, Uyghurs, human rights in China, or indeed can even be instructed to write an intentionally absurd vitriolic screed against the CCP. The same uncensored 27B (dense) will do the same, just at a slower token/s rate.
Similarly there's 'uncensored' variants of Gemma4 31B and other western trained models, which once put through the same process, will also discuss or write just about anything you want, bypassing whatever internal guard rails were attempted in the training data set.
edit: more concerning, and a very legit concern, is that a model is only as good as the sum total of its training dataset, so if something is trained on a steady diet of news sources like Peoples Daily, Xinhuanet and similar in the English language, then it'll have a greater percentage of CCP-approved media publications in its training dataset. No amount of uncensoring it will help with that after the fact.
est 2 hours ago [-]
> China can (and does) use the models to influence the west
OK now that's false information.
You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.
_aavaa_ 16 hours ago [-]
> distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here? ... The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation
Sounds great to me; live by the sword, die by the sword.
eli 8 hours ago [-]
Seems only fair that if LLMs can use copyrighted data for training then they should be able to use cannot-be-copyrighted output of other LLMs.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
mediaman 7 hours ago [-]
This happens all the time. The government can decide legislatively that certain commercial terms are simply unenforceable. Making distillation clauses unenforceable in tort law would be straightforward. They can decide what customers they want to have, but they do not have unfettered rights as to the enforceability of terms governing the relationships between the parties.
eli 6 hours ago [-]
I'm not doubting it's possible to pass such a law, I'm doubting that's it's a practical or worthwhile goal.
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
hnfong 38 minutes ago [-]
You two are talking about different things.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
This is the other thing.
And I think you are both right.
Terr_ 6 hours ago [-]
> Seems only fair
"You're trying to kidnap what I've rightfully stolen!" -- Vizzini
paxys 7 hours ago [-]
It's pretty common to have such laws. OpenAI can put whatever they want in their ToS, but they cannot go back and sue someone for violating those terms if the government has ruled that clause to be unenforceable.
matheusmoreira 5 hours ago [-]
> OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
Correct. It shouldn't be allowed to do that.
jay_kyburz 5 hours ago [-]
Err.. I would like preserve my own right to decide who I'll do business with.
SomeHacker44 4 hours ago [-]
You already do not have that unfettered right in the USA.
CamperBob2 4 hours ago [-]
Then write your own training corpus.
grim_io 7 hours ago [-]
Forbidding distillation is like forbidding using a compiler to make another(perhaps better, more efficient) compiler.
chuckadams 7 hours ago [-]
Lots of software licenses have “non-compete” clauses that forbid you from using it to develop a competing product. Wouldn’t surprise me if there was a compiler or two out there with that restriction, most likely niche languages.
not2b 5 hours ago [-]
It's been common in electronic design automation tools to have license terms like that (forbidding use to create a competing product). However, competing companies have often found workarounds, either by finding loopholes or just breaking rules and hoping not to get caught.
matheusmoreira 5 hours ago [-]
Those clauses should be illegal.
scotty79 6 hours ago [-]
How the hell is non-compete legal in market economy? Competition is one of its core strengths. Why would anyone let anyone opt out of this, even a little bit?
thesmtsolver2 6 hours ago [-]
No country in the world is full free market economy. It is always a spectrum.
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
scotty79 6 hours ago [-]
Chinese companies compete ruthlessly between themselves though. That's how they get this good. Full competition with preventing exploitation by foreign countries seems to be working great for them. American and European protectionism of local rent-seekers can't really compete with that.
cayley_graph 6 hours ago [-]
Yup, fair's fair. Anything else stinks of 'rules for thee but not for me' (a maxim the frontier labs seem worryingly happy to apply, on several counts).
ronsor 8 hours ago [-]
I am immediately sold on this.
Sorry, OpenAI & Anthropic.
magarnicle 4 hours ago [-]
Why would reading copyrighted material ever be an issue anyway? Wouldn't copyright law only apply to what you create and publish using the model? Training on every comic book should already be perfectly legal, as long as you accessed them legally, right? But publishing your own Batman comic using that training is copyright infringement.
What I'm saying is, doesn't the law already cover 1?
_aavaa_ 4 hours ago [-]
Fair use requires more than you accessing the material legally.
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
magarnicle 3 hours ago [-]
> If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
matheusmoreira 5 hours ago [-]
> distillation: why exactly is it bad?
Felony contempt of business model.
qurren 7 hours ago [-]
Government cannot exactly "bar" terms of service. ToS isn't law. The most they can do is say they're unwilling to enforce them.
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
ascorbic 7 hours ago [-]
The government absolutely can pass laws that ban particular contract previsions. They do that all the time. In your analogy for example while they can require you to wear red, they can't require you to be white.
onesociety2022 6 hours ago [-]
Governments can do anything they want by passing a new legislation. In your example, they could easily pass a law that states that any ToS cannot reject service to a customer based on the color of their attire. In the USA, it's obviously already illegal for a business to reject service to a customer based on some protected classes like race.
ButlerianJihad 6 hours ago [-]
The joke is on you! I’m not wearing any attire! Hahaha!
nl 6 hours ago [-]
That's just not true. You can absolutely have terms of service that are illegal, and the government can enforce them.
llm_nerd 8 hours ago [-]
The distillation explanation is classic American exceptionalism: No one could possibly do anything unless they were copying American leaders (where "American" means a bunch of Chinese, Canadian, Europeans and Indians working in the US).
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
villish 1 hours ago [-]
> Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
hnfong 33 minutes ago [-]
> European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
Which Chinese model was it that identified itself as Claude 15% of the time?
llm_nerd 7 hours ago [-]
Models don't have some self identity, beyond what is explicitly handed to them via a system prompt. There have been many, many cases of models identifying as different models by different makers as a basic identity hallucination. They train on enormous volumes of data including lots of people talking about certain makers and models (ChatGPT was actually a super common one given that it became the kleenex of the LLM world). Hence why vendors have to specifically tell it to override that, and if they don't you get lots of funny cases of identity confusion.
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
noncoml 8 hours ago [-]
Don’t know much about how distillation works so please enlighten me here.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
paxys 6 hours ago [-]
You need to do both.
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
numpad0 8 hours ago [-]
known-good prompt-response pairs are more useful than random semi-coherent texts presumably
root_axis 7 hours ago [-]
Because the model can output data in a manner optimized for training a new model, including outputs that were post-trained like RLHF and RLVR.
petilon 8 hours ago [-]
[flagged]
altruios 8 hours ago [-]
This is a silly perspective, inaccurate, and out of bounds framing.
Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.
But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.
But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.
petilon 8 hours ago [-]
If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
lelanthran 34 minutes ago [-]
> If a professor learns from multiple books, generalizes from them and then shares his knowledge he is providing a valuable service. Versus someone who makes a recording of the professor's lectures and resells them to undercut the professor--that guy is not providing a valuable service.
I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?
Where were you going with this?
tux3 8 hours ago [-]
What kind of fresh hell does the sentence "undercut the professor" come from?
Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.
petilon 7 hours ago [-]
We are not really talking about teaching here.
_aavaa_ 7 hours ago [-]
No, we’re talking about an intimate set of tensors, not a human being.
A tree falling and killing someone isn’t tried for manslaughter.
So I don’t care about a hypothetical teacher.
idle_zealot 8 hours ago [-]
What about a student attending lectures of other professors and generalizing what he learns from them, then going on to become a professor?
petilon 7 hours ago [-]
If the student is really good at generalizing we wouldn't even be having this debate because he would've just generalized from the same source materials the professor used.
altruios 7 hours ago [-]
What are you even trying to say: "undercut the professor"...
The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...
Crisco 8 hours ago [-]
No, but the students that learn and distill what the professor teaches are not obligated to use that information only how the professor wants them to.
petilon 8 hours ago [-]
Can the professor refuse to teach some students, or must he teach all comers?
xboxnolifes 8 hours ago [-]
One can teach whoever they want to or don't want to. If one joins a university, that changes things. They are now part of an organization larger than themself.
lostmsu 8 hours ago [-]
glhf under the circumstances
smarf 8 hours ago [-]
'why is reselling stolen stuff bad'
if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.
mywittyname 7 hours ago [-]
More like, is a professor who learned from books prohibited from writing his own books on the subject?
petilon 7 hours ago [-]
He is prohibited from regurgitating source material, of course! But if he generalized from the books he read and really learned the subject--and even made new connections between ideas--then he is free to write his own book.
mywittyname 7 hours ago [-]
He is not prohibited from "regurgitating source material" in many cases. Facts are free. It doesn't matter who first measured Young's Modulus of aluminum, anyone may state that fact as originally presented.
The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.
8 hours ago [-]
janalsncm 7 hours ago [-]
1) No one is asking Anthropic to give tokens for free, but at market rates.
2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.
_aavaa_ 7 hours ago [-]
Who cares, a LLM isn’t a person.
bluegatty 6 hours ago [-]
Making an LLM from raw data is value-add.
Distillation is just value extract.
It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.
I think we start by recognizing that ... and then try to figure it out from there.
'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
nemomarx 6 hours ago [-]
What makes the Internet raw data in a different way? wasn't it mostly worked on by people first?
bluegatty 6 hours ago [-]
There is value add in AI irrespective of how the data got to what it is.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
nemomarx 6 hours ago [-]
Okay, so if the chinese models are used everyday, do they become a value add? Like what's the line you're drawing here. Amount of value it creates?
bluegatty 4 hours ago [-]
Designing and creating an LLM from nothing is a monumental feat of Engineering and 'value add'.
Copying something is not.
Programming Microsoft Word is value add, copying the code is not.
Copying design ... there are some question marks there.
It's extremely easy to understand at it's core.
What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.
"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"
The training data used is part of all of this is a separate but related question.
scotty79 6 hours ago [-]
> Making an LLM from raw data is value-add.
> Distillation is just value extract.
There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
bluegatty 6 hours ago [-]
I think that's kind of fair, but it still fits within the context of 'some things are value add' and 'more or less than others'.
We ought to identify that and integrate that into our thinking.
ArtRichards 15 minutes ago [-]
New, smaller models can outperform the previous generation's foundation models.
What if there's a way to extract the commodity of intelligence from smaller models?
I've seen for many use cases it's well enough. :)
OleksandrC 8 hours ago [-]
The article makes a point about agent harnesses being sticky (the supposed moat). I have been building my own agent harness for a while, and I can tell with confidence that the harness almost does not matter, the entirety of the AI magic is the model itself. The harness can be almost barebones (like, for example, mini-swe-agent used for benchmarks), and yet the model still does the task just fine.
So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
bze12 7 hours ago [-]
By the harness I believe he means the entire end-user product experience, not specifically the harness code. I’ve mostly stuck with codex because their Mac app is better and I’ve gotten used to running automations through it. The more workflows they can build around this (design tools, collaboration, etc), the better chance of lock-in.
> If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
Sol- 8 hours ago [-]
For me, harnesses are mostly sticky insofar as the model providers only allow you to use their subsidized plans through their own harnesses, unfortunately. But of course switching model + harness is an option.
dansquizsoft 5 hours ago [-]
Facts, I was able to code a personal self improving harness in a weekend (something a bit more similar to Hermes or OpenClaw at the time but with a more expansive set of features for my use cases and requirements) and it works great for 90% of the tasks I would use Claude Code or Codex (now ChatGPT App) for, with the remaining 10% being able to be implemented with a few more prompts from within the harness itself.
For this reason alone I would also argue that the idea about an agent harness being sticky is a non-starter long-term.
hdz 8 hours ago [-]
The harnesses will tend towards commoditization, but for now the harness quality matters a lot. Especially for non terminal harnesses.
neutronicus 5 hours ago [-]
Yeah it certainly feels like the harnesses are pretty minimal value add on the token pipe
jke_kang 8 hours ago [-]
People seem to conflate "made in China" with "can't be trusted." id argue the bigger distinction is open vs. closed. An open model can be audited, fine-tuned, and technically run entirely on your own hardware. A closed model is basically "trust us."
mrinterweb 7 hours ago [-]
Open weight models are much more auditable than closed models, but could still hide backdoors that could be near impossible to detect.
chrsw 7 hours ago [-]
Correct. We need open weights, open code and open data. If nobody else can reproduce what someone did there will always be security questions. Even if we can reproduce it there could still be security concerns but it's more realistic to investigate yourself.
nl 6 hours ago [-]
I'm all for open models, but people seem to misunderstand what they are. They aren't the same thing as open source code!
> open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
chrsw 6 hours ago [-]
You don't want to replicate the exact model, you want to build a system of similar capabilities.
nl 4 hours ago [-]
Great, but that seems a different concern to the auditability of a model.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
Ericson2314 3 hours ago [-]
Deterministic seed
essentia0 7 hours ago [-]
Exactly what are the possible 'security issues' of self hosting an open weights model?
perching_aix 7 hours ago [-]
It may have been backdoored during training, potentially causing it to randomly start wreaking havoc at runtime, possibly in a clandestine manner (e.g. sneaking in bugs into generated code).
urams 4 hours ago [-]
> We need open weights, open code and open data.
Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.
wyrdcurt 6 hours ago [-]
In my opinion, the big issue with that argument is that advances in interpretability research and steering conceivably could, and probably will, render moot that (as of now, purely hypothetical) risk of subtle sabotage for open-weight models... but not for closed models.
_factor 5 hours ago [-]
It’s not hypothetical. Magic strings are a known and implemented feature for standard model interaction. Nearly impossible to detect unless you know where to look with current technology.
wyrdcurt 3 hours ago [-]
Maybe I should clarify. As I understand it, the kind of vulnerability being discussed is something like a Chinese model invisibly "realizing" that it's working on an American project, and then deliberately leaving subtle security bugs in its generated code for Chinese hackers to later exploit. As far as I know, that scenario is hypothetically possible, but has never been demonstrated to happen in the wild. Admittedly, I could be wrong about that! If anyone has evidence to the contrary, I'd love to see it.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
CamperBob2 4 hours ago [-]
As long as I can say, "Model A, look for security holes in this code by Model B," I don't see this being a serious problem.
It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.
galacticaactual 5 hours ago [-]
Oh really. How'd that work out for security in open source.
weird-eye-issue 3 hours ago [-]
I think that perception of China has been shifting and will look quite different over the next few years
credit_guy 4 hours ago [-]
People who claim that the Chinese open weight models have some type of manifest advantage don't realize that the close weight models have a huge advantage as well: the researchers from OpenAI, Anthropic, Google, xAI, Meta are not dumb, they can read the white papers written by DeepSeek, Moonshot, etc, and they can inspect all those architectures and they can pick and choose the best tricks there are out there, and of course, they have access to their own in-house secret sauces.
Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.
But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:
1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.
2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.
3. The frontier labs are also investing more and more in building an ecosystem around their models.
I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.
People who claim that Postgres has some type of manifest advantage don't realize that Oracle has a huge advantage as well…etc
killingtime74 2 hours ago [-]
If the Google and meta engineers are not dumb how come they consistently trail behind the frontier labs and even the Chinese labs with a fraction of the funding.
Probably bad leadership
hnfong 20 minutes ago [-]
I always suspect they have the most to lose if legal decisions on copyright issues don't go their way.
Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.
This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.
And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.
Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.
(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)
kubb 1 hours ago [-]
That plus they don’t distill so they have worse RL examples.
protocolture 3 hours ago [-]
>they have access to their own in-house secret sauces.
I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.
If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.
>3. The frontier labs are also investing more and more in building an ecosystem around their models.
Theres nothing there that isnt immediately replaceable.
eeiei 4 hours ago [-]
It’s giving desperate!
bg24 6 hours ago [-]
I think in general rest of the world needs to take notice (not saying afraid), starting with the US. It cannot be taken for granted that China's frontier labs will be a few months behind. They might be at par or exceed.
The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.
At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.
Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
awakeasleep 4 hours ago [-]
In the earlier days of the USA we did the same thing, with our government having an industrial policy that fed US industry and put us ahead of Great Britain.
It doesn't have anything to do with the form of government, it has to do with the aims of the government.
Barrin92 3 hours ago [-]
What is this panic and protectionism supposed to be good for? This is open source software, there is no "AI industry", there's virtually nobody employed in this. "Domestic AI" makes about as much sense as a domestic Linux kernel. If the Chinese want to subsidize the world's water and energy use to supply the world with chatbots good luck to them. There's no need for guardrails or fast track data centers, they can plaster their entire country with data centers to churn out slop, I'm glad we don't
hnfong 28 minutes ago [-]
Only those who intended to use the new technology as a tool to achieve dominance and control are freaking out.
And it's the only viable tool the US has left. It's reasonable they are freaking out.
bg24 3 hours ago [-]
I think it is deeper than that. "LLM => peak of productivity" takes way less time and effort than "Linux kernel => any productive work". Compare Dec 2025 vs July 2026 models in terms of capabilities.
Nobody can predict 5 year out. However, the country that can be ultra efficient by making their governance, health, manufacturing, military, etc AI-native will be far ahead in the game.
Barrin92 2 hours ago [-]
There's no evidence that these things make anyone more productive, in fact the opposite, empirical research suggests users are less efficient while thinking they're more efficient. You're invoking "AI" the same way the blockchain people did before they all finally died out. "You need to put the government on the blockchain bro, you're gonna lose out to the future bro", turns out you don't
But that's not really even relevant to the debate. Insofar as software has made the world more efficient, it doesn't matter where it's written. That's the point of open source software, there's no tacit knowledge. When a country loses its nuclear engineering capacity that's dangerous because it takes a long time to rebuild. The only reason the Chinese are already competitive on LLMs but haven't managed to make a state-of-the-art jet engine is because the latter, unlike software, is difficult to copy and you don't need to worry about something you can copy.
eeiei 3 hours ago [-]
Ok…
0x38B 3 hours ago [-]
Excellent article; the argument towards the end for allowing distillation for US companies is compelling:
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
protocolture 3 hours ago [-]
Very good point.
That would prevent the facebook strategy of sucking up MySpace users and then defending TOS that prevent other social media apps from doing the same to them.
throwa356262 8 hours ago [-]
According to openAI's own @deanwball:
Even OpenAI isn't buying this distillation talk:
I’m struggling to understand this perspective. Is he using the words accelerationist/decelerationist in a sense other than the obvious one?
EDIT: I searched his twitter history and discovered that his argument is basically “if you drive down costs, then OpenAI will have less money to invest in development, slowing down the overall rate of AI progress.” IMO this take betrays an overwhelmingly stupid degree of exceptionalism, but I guess that’s what I’d expect from someone working at OpenAI.
aesthesia 3 hours ago [-]
Nah, the point is that if models are commoditized and there's no hope of making significant profits, no one is going to be willing to make the massive investments necessary to continue pushing the scale frontier. How large a training run do you expect investors to fund out of the goodness of their hearts?
twelvedogs 54 minutes ago [-]
what good does it do to invest in a solved problem?
if open models are good enough then it doesn't matter, if they aren't then there's probably a return available in investing there.
protocolture 3 hours ago [-]
The real problem is that I could probably solve even biggerer issues if the investment went to me instead of OpenAI, so OpenAI has a moral imperative to shut down and send me all their money right now.
eeiei 4 hours ago [-]
What a delusional f-wit lol he wants protection of profits for reinvestment?
Every company wants that!
3 hours ago [-]
nothercastle 8 hours ago [-]
This guy is predicting AI covid escaping from a Chinese lab. I find that kind of silly
an0malous 5 hours ago [-]
So he thinks open weight models will lead to “AI communism” and “dystopian hell” and in the very next point proposes that the US create a federal agency to discourage the use of open weight Chinese models. The motivated reasoning in this post is unreal.
Can you or someone please explain several of the claims made in this tweet?
"I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. I don't understand this at all.
Can someone in the know please use plain layman's terms to explain what this tweet is about?
slopinthebag 7 hours ago [-]
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.
Barrin92 7 hours ago [-]
>Can someone in the know please use plain layman's terms to explain what this tweet is about?
The Silicon Valley people like this openai guy, high on their own supply, are convinced they are building some machine god that will either bring about the end of the human race or utopia, they therefore cannot understand why the Chinese (or any other normal person on earth) are not afraid of chatbots and have other things on their minds.
eeiei 4 hours ago [-]
[dead]
paxys 6 hours ago [-]
> "I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I assume they mean the risk of opening up "forbidden" knowledge to the masses without adequate control, which the CCP hasn't historically been known to do.
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Yann Lecun is a pioneer in the field of AI and Meta's former AI head. He is famously anti-LLM, and considers the entire technology a dead end to achieving human-level AI. The author is saying the CCP has similar views (that LLMs aren't going to get exponentially better/lead to AGI) which is leading them to not control these models as tightly as they otherwise would.
> Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
"AI accelerationists" = people who want AI to progress. According to the author these people should not celebrate open models because open source = less commerical value in LLMs = less investment into the field (because how are companies going to get returns?), and this will ultimately lead to slower growth.
The last bit is about government controlling AI vs commercial companies. According to the author the former is a dystopian hellscape.
IMO even if you think his points make sense, his job title ("head of strategic futures @openai") means they should all be taken with a massive grain of salt.
nothercastle 8 hours ago [-]
Chinese ai is bad. It’s slowing down progress and it’s so bad we called out the c word and asked for more regulation. Basically advocating for more government assistance to openai
spenvo 5 hours ago [-]
"Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence"
That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
c0decracker 5 hours ago [-]
But.. if you are running Chinese model in the US, what difference does it make? Isn't the whole "scare" (khm khm) with Kimis is that now I don't need Claude, cause I can run Kimi on my own hardware in my own datacenter and it's maybe not as good as Claude July edition but it's is as good as Claude January edition.
richardlblair 4 hours ago [-]
It doesn't need to be as good. You can route to the appropriate model and save so much money.
I have sonnet do the thinking, deepseek does all the tasks. I've massively reduced costs with this approach.
spenvo 5 hours ago [-]
Sure, and I think that flexibility further undercuts his "frontier labs will be fine" take, which depends on top US labs having pricing power.
deaux 4 hours ago [-]
> This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?
This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.
And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
Alien1Being 9 minutes ago [-]
The options are to use LLMs from a country run by a psychopathic regime or alternatively to use Chinese LLMs.
zkmon 4 hours ago [-]
What's wrong if the roles of USA and China are reversed in technology? Why does the rest of the world care? It's not as if USA has done a great good for the world, and China has evil intentions towards the world. Infact it is the opposite in the case of AI so far.
mattas 3 hours ago [-]
"Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute."
It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.
The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
simonreiff 6 hours ago [-]
I fully agree with everything in this essay. Make distillation fair use. And let us use Mythos/Fable and Sol and successor or future models for all cybersecurity purposes.
softwaredoug 4 hours ago [-]
Haven’t we been in this “China is 3-6 months behind” for a while now (maybe up to a year? Longer?)
The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story
Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.
It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
nunez 54 minutes ago [-]
I really enjoyed reading this.
This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.
Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)
I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.
Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
overfeed 4 hours ago [-]
> [Anthropic/OpenAI] are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs. Second, intelligence isn’t in fact a perfect commodity, in part because applied intelligence makes itself smarter
Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
golly_ned 4 hours ago [-]
> I expect the inference market to grow much faster than training costs
This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.
But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
softwaredoug 7 hours ago [-]
> By the same token, don’t expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it’s just as much of a mistake to pretend like distillation doesn’t give Chinese labs a big advantage.
I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].
If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.
A company making a decision to allow use of chinese models is a company also choosing to send tons of various credentials to chinese model companies. These will just get scooped up, OpenAI and Anthropic can probably hack into anything at this point if they wanted to.
Tostino 1 hours ago [-]
But, they are releasing the weights very shortly (or already have for some of the models discussed). For a very large company, you can purchase or rent the hardware yourself to serve the models.
Or any US hyperscaler with GPUs to spare can decide to serve the models for a reasonable cost/token.
You don't have to send China your data.
ab_wahab01 2 hours ago [-]
Honestly, as someone from a developing country, this shift is good for us. US frontier models are too expensive for us to use regularly. Chinese open-source models/subscriptions are really good to use.
minraws 8 hours ago [-]
Me I am, so very afraid of actually decently priced inference.
ggm 6 hours ago [-]
A reminder any comment about risk FROM china, invites a "Tu Qoque" facing the other way. The paranoia here is probably fully symmetrical.
I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.
A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
coretx 4 hours ago [-]
The best model is the model that runs best on your hardware.
hexator 7 hours ago [-]
I'm worried that any ban on Chinese AI models might be an excuse to get mass surveillance.
onesociety2022 6 hours ago [-]
You don't need mass surveillance to enforce such a ban. Once the US Govt declares Chinese AI models are banned, no US business will use them nor distribute them. Any cloud service that rents out GPUs in the USA will explicitly prohibit the use of Chinese open model weights in their terms of service (you open yourself to a lawsuit if you violate their ToS). Any Tokens-as-a-Service provider will refuse to serve those tokens to customers in the US.
Sure as an indie hacker, you could go download the weights for a Chinese model with a VPN, and then attempt to run it at home by building your own GPU cluster but these large models require quite expensive hardware to run on and so it makes it less likely than anyone would invest that much capital to do something that is illegal. There's no way for them to sell a legal service using those tokens. So it can only be strictly for personal use (the Govt won't care because very few people will have that kind of money and risk appetite). The other option will be that there will be some shady third-party providers in foreign countries who are willing to sell tokens from these models to US consumers knowingly.
chockablock 4 hours ago [-]
> Any Tokens-as-a-Service provider will refuse to serve those tokens to customers in the US.
So under a ban rest-of-world gets to use cheap open-weight models but American companies/individuals must only use only ‘approved models from US for-profits’? Doesn’t seem like that kind of protectionism will be popular or politically tenable. Not so long ago US chose cheap TVs over maintaining the country’s manufacturing base.
(Despite what you wrote it’s also really hard to imagine that enforcement wouldn’t leak like a sieve. Unser sufficient economic incentives [which are the predicate for the ban], loopholes will be found.)
nl 6 hours ago [-]
> because U.S. open weight model makers must follow the frontier labs’ terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn’t it be better if western open weight model makers could go to the source?
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
But it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum.
I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.
The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.
IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
alizaki 6 hours ago [-]
There is no “Chinese LLM”. Each “lab” is distinct and their models behavior is as unique as those from OpenAI and Anthropic
wmf 6 hours ago [-]
Somehow a certain set of labs are all releasing open weights and a certain other set of labs are closed weights.
dofm 5 hours ago [-]
Somehow the two main closed weights frontier models come from two companies with HQs about two miles apart, and the CEO of one used to work for the other.
purplepatrick 4 hours ago [-]
Commenting wholesale on some folks who are asking for hard evidence. I cannot provide that either but can contribute some empirical data.
I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.
As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.
Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.
I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.
So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.
Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
josht 5 hours ago [-]
Someone (anyone!) get David Sacks on the horn and tell him to read this.
Havoc 7 hours ago [-]
oh wow - hadn't realized they decided to opensource Qwen 3.8 Max. That's pretty big news.
fellowniusmonk 8 hours ago [-]
The U.S. "executive" class is so obsessed with the "exploit" part of the explore/exploit cycle that it's very clear they are prematurely closing advancement. Better a little money and power for them now than a lot of money and power for their country/humanity.
This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.
You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
XorNot 8 hours ago [-]
My personal hypothesis here is the Chinese government looked at the game and simply decided not to play:
An astute Chinese analyst could reasonably forecast that they had little chance of controlling the AI market due to sovereign trust issues, but would also note that AIs are just software.
When the dust settles the US still won't have factories, and the real value of AI models is still going to be embodying them and getting them to do real, consumer facing work.
Perhaps the most striking thing about the AI boom is how quickly the US abandoned the veneer of local manufacturing in favor of more expensive buildings producing nothing you couldn't make anywhere else on the planet...from imported parts.
_carbyau_ 6 hours ago [-]
Yeah, how much of this is China waving distracting AI hands over here while the US Genius-In-Charge watches and completely ignores reality.
NooneAtAll3 8 hours ago [-]
I don't understand the premise in the beginning
how is running servers supposed to be 0 cost, while running ai inferrence isn't?
cheema33 6 hours ago [-]
> how is running servers supposed to be 0 cost, while running ai inferrence isn't?
For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue.
If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher.
Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way.
throwawayffffas 8 hours ago [-]
A typical server that costs 10k to 30k to own and operate can serve between hundreds and thousands of requests per second of a traditional web application like facebook for 2-4 kW of power, the marginal cost of each request is effectively zero.
A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.
pupskipper 5 hours ago [-]
The fact that Anthropic has a model like Mythos means that counterpart countries like Russia and China are not far behind, if they haven't already developed something similar or better.
3 hours ago [-]
anuramat 5 hours ago [-]
> Russia
lmao
npn 1 hours ago [-]
what a horrible article. full of misinformation and dishonesty.
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is.
but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency.
and they even plan to build their own inference hardware, too.
and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
and you say there is nothing to afraid?
jmclnx 8 hours ago [-]
One thing I have not seen mentioned between Chinese AI vs US, population.
China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.
Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.
So I believe, China will end up owing AI.
gerdesj 8 hours ago [-]
"China will end up owing (sic) AI"
I think you hit the nail on the head - right there!
7 hours ago [-]
zzzeek 4 hours ago [-]
the leader of China praised Open Source in a speech. Crazy times
panchtatvam 2 hours ago [-]
Is this article written by AI ? It looks so.
kdqed 2 hours ago [-]
I'm honestly more afraid of Claude
lenerdenator 3 hours ago [-]
It really is amazing that China went from the country that hacked Google out of its market to a trusted source of AI in the tech world.
Joel_Mckay 3 hours ago [-]
Distilling models using more advanced LLM is not a new phenomena. It is a cost effective strategy in a highly competitive emerging field.
Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
icase 2 hours ago [-]
not enough people
alfiedotwtf 4 hours ago [-]
Answer: US investors
jdw64 8 hours ago [-]
While intelligence is said to be a replaceable commodity, oil and copper can be used in nearly the same way even if you change suppliers as long as the quality grade is matched. However, I question whether two models that produce the same benchmark answers are actually interchangeable in real world use.
Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
zuzululu 5 hours ago [-]
My thinking is that with the current narratives out of washington we are on track for a ban on Chinese models and possibly sanctions against Chinese AI companies
I think it is the right move to protect American interests
marwaneet 4 hours ago [-]
i think most is vcs
kmeisthax 2 hours ago [-]
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.
Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.
And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."
[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".
[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
6 hours ago [-]
chews 8 hours ago [-]
later secondaries investors in openai/anthropic. It's like time traveling into the spacex ipo.
sjreese 6 hours ago [-]
Kellogg School of Business -- he said -- token as a commodity and therefore Open AI is constrained .. ha ha ha hee hee ha .. Well... you build a better mousetrap, and DeepSeek, K3, and ByteDance are just that -- just as good and fit to purpose -- What is needed is to build on top of -- not paniteir (invade privacy and kill people with the information) -- not USMC AI -- use PI's as overwatch killer drones -- but how can I make harder steel, longer-lasting, seawater-resistant concrete, faster time to build housing, better enforcement of USDA rules and FDA adverse enforcement, and better EPA water cleanup, a better FTC for consumer goods -- that is, if I buy an item, that item is safe and built to purpose -- ANYONE not talking about public protection of consumer rights usng AI, is wasting your time
magarnicle 4 hours ago [-]
Has someone replaced your return key with a double-dash key?
andrewdubinsky 44 minutes ago [-]
[dead]
nttylock 3 hours ago [-]
[flagged]
ngl999 4 hours ago [-]
[dead]
brandopn 6 hours ago [-]
[dead]
warshinder 3 hours ago [-]
People with 401k’s and retirees?
sharadov 7 hours ago [-]
What makes the Chinese models this good? I don't believe it's distillation alone.
This from OpenAi's Head of Strategic Futures
"Some observations on Kimi:
It's a very good model! I don't think its performance can be explained away by distillation or anything like that"
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
To win on the AI front by any means necessary.
nl 6 hours ago [-]
> China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
Havoc 7 hours ago [-]
>What makes the Chinese models this good?
Why wouldn't it be? China is pumping out AI research and researchers at a staggering pace and there is no inherent reason why western models should be better
nnm 1 hours ago [-]
Looks like authored by AI. With quite some reasoning, but no real data to back up main point.
oeatwell 54 minutes ago [-]
I think frontier labs should start building ecosystems by partnering with companies that already have software products, and even collaborating with hardware manufacturers. The ultimate goal should be to create a much broader range of products that integrate naturally into people's everyday lives.
The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.
samtp 14 minutes ago [-]
This is a really ironic comment given how the US gov is getting politically involved in pretty much all science research funding and speech at the moment.
Aperocky 36 minutes ago [-]
> humanities, creativity, entertainment, or financial transparency
> monetizing these strengths
> real advantage over China is freedom
Please tell me if I'm unfairly paraphrasing but these seem to be your main argument and they seem to be oxymorons
1. US frontier lab unit economics are better 2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.
For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.
For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.
And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.
And every improvement in model capability makes it increasingly easier to make your own tools.
Wrote more on this in a blog post that has an earlier HN discussion: https://news.ycombinator.com/item?id=48982061
Direct link: https://larrysalibra.com/ben-thompson-is-wrong-us-frontier-l...
What is the cost of AI? The single largest ingredient is Nvidia profit margin.
Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.
Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.
Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)
What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.
I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.
Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.
Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)
Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy
Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)
Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.
The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.
Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.
All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.
Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.
This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.
[1] https://www.iea.org/reports/electricity-2026/prices
If I had to wager why, I'd say it's due to embracing solar on massive scales recently. Only a few years ago the US was competitively priced.
[0] https://www.globalpetrolprices.com/compare_countries/USA/Chi...
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]
I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.
It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?
I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.
Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.
If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.
But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.
And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.
> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?
Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).
The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.
Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.
> This is because US labs are leading on cost efficacy of inference ($/task)
It’s possible, but I would need to see better data on this.
>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.
IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.
So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.
Over time, this compounds. More profits means more investments. More market share means more control.
I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.
As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.
I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.
This is the story for Nvidia/AMD or cloud providers rather than OpenAI.
> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
It seems like there would be problems with this on both ends.
For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.
Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.
And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.
Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.
I notice that the article, and this discussion, hasn't mentioned or considered local models.
We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.
If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?
This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.
But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.
Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.
The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.
From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.
But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.
That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)
It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".
Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.
[0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).
But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.
It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.
I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.
Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.
Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.
But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.
* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.
* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.
* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].
OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.
[0] https://www.mindstudio.ai/blog/anthropic-inference-margins-7...
[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware" https://x.com/rohanpaul_ai/status/2079027313455550839
I would say the weirdest thing about the article was that it was entirely in terms of business value. But commercial businesses are at best just a stop-gap measure, to a Marxist.
Using a Chinese LLM will not put a Marxist under your bed.
Did you know Gemini is shockingly bad at French poetry? Hasn’t stopped me for using it for all other tasks though.
Interesting tangent - it might be taught to introduce stealthy backdoors in your company though. Maybe even across multiple PRs where each session puts a small chink in the armor, and together they allow unlimited access to the attacker who knows about them.
After all, LLMs are mostly black boxes. How comfortable would you be running a Chinese compiler?
Its the user base (with ads and upselling) and proprietary wrappers which will make money for typical customer.
Even enterprise customers arent going to be spending a lot on tokens. Once labs no longer have to subsidize trainings tokens costs will drop 10x and once models get burned on chips costs will drop 10x more and you physically won't be able to burn significant number of tokens unless you're deliberately trying to.
https://epoch.ai/data/ai-data-centers
They’re going to try their best to offload these investments into our pensions before the inevitable crash.
https://finance.yahoo.com/markets/stocks/articles/goldman-sa...
Too many people here buy high and sell low.
When Goldman makes statement, half the time it is prepped by an associate or two that has minimal experience and goes through a MD that enjoys the gloom and doom. That’s why they publish. Goldman makes money on both sides of a trade.
Correct. These chinese labs has proven that having just the model is not a moat, and the safety concerns were all just attempts at regulatory capture.
This is why labs like OpenAI and Anthropic are panicking and are racing to the exit before their valuations start being questioned.
for VCs, breaking even is losing
https://github.com/JustVugg/colibri
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
For companies, these decisions are very sticky. Companies go through a lot of red tape to get anything purchased and approved, then they discourage change because it's a lot of work.
So the product that gets a foothold in a company sticks for a long time.
Then a couple years later a sales person convinces an exec that they can save some money by switching, so the switching game begins. Not necessarily motivated by the better product, mostly the price. My wife's company keeps switching their tools out from under everyone every year or two. Just when they get everything stabilized and everyone familiar with the new tool, some new contract is signed that moves them all to some other company's suite.
Similarly, I have made no ground in arguing to try to get Codex at the company I work for, which got Claude Code a year ago and sees no reason to go through the whole process of setting up any alternatives when Claude Code already works and is at the frontier.
You can choose a selection of different models within it, but you're not using Codex or Claude Code.
They communicate through my own harness, and it's working pretty well so far. claude code is being overtaken by codex however because I noticed lately the accuracy of the latter is the best.
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
its so distracting seeing these types of confision.
every plugin is already just multimodaling their targets.
Probably the only reasons I would seek change are economical.
“I don’t really have a strong preference between the two” is another way of saying “the product isn’t sticky”, which is another way of saying “this provider has very little room to increase margins”
There’s little difference between Coke and Pepsi and the barrier to switching is nil, yet clearly the products have stickiness. People have slight preferences and become familiar with the brand and then engagement becomes habitual.
The effects on margins are irrelevant to this.
The cost of me moving around these different AI models and harnesses was pretty much 0.
Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question
https://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...
Genuine question: generalized up from individual models to “models from country X”, is there any country that doesn’t have this exact risk?
In China, there is 1 party. 1 view. 1 definition of the Truth.
This is just oriental despotism paranoia, whatever cutsie repetition slogan you come up with is not a serious argument.
I asked Grok "Tell me about the gaza genocide" and it write a IMHO balanced answer comparing why genocide is and isn't the right term. [0]
ChatGPT 5.6 Sol only explained why people call it a genocide and did not go in as in depth as Grok did for why people don't agree with the term. [1]
The only unsaid response (to me) here is the model should have declared that it was not a genocide, and because these models explain why it was a genocide, they are bad?
[0] - https://grok.com/share/c2hhcmQtMi1jb3B5_c078bc15-ef5e-42bc-b... [1] - https://chatgpt.com/share/6a5ef005-0148-83ec-828d-65ae7a7f42...
In international law, that is the ultimate authority of what is/isn't a genocide. It's objectively, legally speaking, a genocide.
1/ I don't see in the responses where the model says it is or isn't a genocide. Can you share the snippet from each, I included the logs above?
2/ I can't find a source on the UN ruling that you mentioned. I am not interested in the findings of an investigative body, just the official UN ruling. Can you share? ChatGPT (and myself) can only find this [0], which is a second round of written submissions.
[0] https://www.icj-cij.org/node/206406?utm_source=chatgpt.com
[1]: https://www.aa.com.tr/en/science-technology/xai-s-grok-tempo... [2]: https://www.ohchr.org/en/press-releases/2025/09/israel-has-c...
The UN ruling is discussing in the South Africa vs Israel case [0], which has not been ruled on yet.
[0] - https://www.icj-cij.org/node/206406
There are also half a dozen other companies from China continuously hammering our clients’ websites.
I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.
Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region
credit:
'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc
Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.
Assuming that China only distills is a huge mistake.
It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.
Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.
The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.
Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.
This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....
There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.
2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.
Meanwhile if you are in the US, DHS has already subpoenaed social media sites looking for people who made anti-ICE posts and I can't imagine they consider subpoenaing AI conversations off limits https://www.nytimes.com/2026/02/13/technology/dhs-anti-ice-s...
One of the interesting things is that through a fairly rudimentary process which is being done by 3rd party amateurs who've downloaded the open models, models like Qwen 3.6 35B-A3B (or 27B) can be fully 'uncensored' when turned into GGUF files.
I have an uncensored Q8 version of Qwen 3.6 35B-A3B here that will very happily output information about Tiananmen Square, Uyghurs, human rights in China, or indeed can even be instructed to write an intentionally absurd vitriolic screed against the CCP. The same uncensored 27B (dense) will do the same, just at a slower token/s rate.
Similarly there's 'uncensored' variants of Gemma4 31B and other western trained models, which once put through the same process, will also discuss or write just about anything you want, bypassing whatever internal guard rails were attempted in the training data set.
edit: more concerning, and a very legit concern, is that a model is only as good as the sum total of its training dataset, so if something is trained on a steady diet of news sources like Peoples Daily, Xinhuanet and similar in the English language, then it'll have a greater percentage of CCP-approved media publications in its training dataset. No amount of uncensoring it will help with that after the fact.
OK now that's false information.
You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.
Sounds great to me; live by the sword, die by the sword.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
This is the other thing.
And I think you are both right.
"You're trying to kidnap what I've rightfully stolen!" -- Vizzini
Correct. It shouldn't be allowed to do that.
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
Sorry, OpenAI & Anthropic.
What I'm saying is, doesn't the law already cover 1?
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
Felony contempt of business model.
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
https://m.economictimes.com/industry/renewables/china-wto-co...
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
Public libraries, in this instance, is curated data from all the internet, obtained through not legal means (I don't have a problem with this other than lack of attribution, being copy-left). Just to be clear.
But in answer to your incredibly leading and inaccurate framing... they are required (by their job title) to teach to those who who show up in the classroom, it's not their place to discriminate against anyone/thing (even those like itself (other robots)) that also show up in the classroom.
But you can't teach at a university using only knowledge learned from the library. you need a degree. You are free to teach at the park, where anyone can hear you. public in -> public out.
I'm confused now; isn't the LLM that trains on that professor's lectures, videos and textbooks undercutting him?
Where were you going with this?
Teaching isn't a race to the bottom. You don't undercut teaching by giving more lessons, just like you don't slight the hospital by performing CPR.
A tree falling and killing someone isn’t tried for manslaughter.
So I don’t care about a hypothetical teacher.
The further you try to constrain this topic into this illformed analogy the weirder it becomes. If we start off with a better analogy...
if the professor took all human knowledge, much of which was explicitly not free, and used it to make a for-profit knowledge machine that extrudes unreliable summaries of that knowledge, then yes, being obligated to teach for free would be a fitting punishment.
The professor is free to lift all the facts and formula they want. They just need to rephrase explanations. Which is pretty much what an LLM is going to do.
2) Any professor who tried to ban students from posting lecture notes online would be immediately mocked.
Distillation is just value extract.
It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.
I think we start by recognizing that ... and then try to figure it out from there.
'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
Copying something is not.
Programming Microsoft Word is value add, copying the code is not.
Copying design ... there are some question marks there.
It's extremely easy to understand at it's core.
What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.
"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"
The training data used is part of all of this is a separate but related question.
There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
We ought to identify that and integrate that into our thinking.
What if there's a way to extract the commodity of intelligence from smaller models?
I've seen for many use cases it's well enough. :)
So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
He talks about this in another recent essay https://stratechery.com/2026/anthropics-safety-superpower/
> If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
For this reason alone I would also argue that the idea about an agent harness being sticky is a non-starter long-term.
> open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.
Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.
But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:
1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.
2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.
3. The frontier labs are also investing more and more in building an ecosystem around their models.
I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.
[1] https://www.anthropic.com/careers/jobs
Probably bad leadership
Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.
This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.
And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.
Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.
(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)
I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.
If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.
>3. The frontier labs are also investing more and more in building an ecosystem around their models.
Theres nothing there that isnt immediately replaceable.
The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.
At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.
Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
It doesn't have anything to do with the form of government, it has to do with the aims of the government.
And it's the only viable tool the US has left. It's reasonable they are freaking out.
Nobody can predict 5 year out. However, the country that can be ultra efficient by making their governance, health, manufacturing, military, etc AI-native will be far ahead in the game.
But that's not really even relevant to the debate. Insofar as software has made the world more efficient, it doesn't matter where it's written. That's the point of open source software, there's no tacit knowledge. When a country loses its nuclear engineering capacity that's dangerous because it takes a long time to rebuild. The only reason the Chinese are already competitive on LLMs but haven't managed to make a state-of-the-art jet engine is because the latter, unlike software, is difficult to copy and you don't need to worry about something you can copy.
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
That would prevent the facebook strategy of sucking up MySpace users and then defending TOS that prevent other social media apps from doing the same to them.
https://xcancel.com/deanwball/status/2078133895766114412#m
I’m struggling to understand this perspective. Is he using the words accelerationist/decelerationist in a sense other than the obvious one?
EDIT: I searched his twitter history and discovered that his argument is basically “if you drive down costs, then OpenAI will have less money to invest in development, slowing down the overall rate of AI progress.” IMO this take betrays an overwhelmingly stupid degree of exceptionalism, but I guess that’s what I’d expect from someone working at OpenAI.
if open models are good enough then it doesn't matter, if they aren't then there's probably a return available in investing there.
Every company wants that!
"I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. I don't understand this at all.
Can someone in the know please use plain layman's terms to explain what this tweet is about?
I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.
The Silicon Valley people like this openai guy, high on their own supply, are convinced they are building some machine god that will either bring about the end of the human race or utopia, they therefore cannot understand why the Chinese (or any other normal person on earth) are not afraid of chatbots and have other things on their minds.
I assume they mean the risk of opening up "forbidden" knowledge to the masses without adequate control, which the CCP hasn't historically been known to do.
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Yann Lecun is a pioneer in the field of AI and Meta's former AI head. He is famously anti-LLM, and considers the entire technology a dead end to achieving human-level AI. The author is saying the CCP has similar views (that LLMs aren't going to get exponentially better/lead to AGI) which is leading them to not control these models as tightly as they otherwise would.
> Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
"AI accelerationists" = people who want AI to progress. According to the author these people should not celebrate open models because open source = less commerical value in LLMs = less investment into the field (because how are companies going to get returns?), and this will ultimately lead to slower growth.
The last bit is about government controlling AI vs commercial companies. According to the author the former is a dystopian hellscape.
IMO even if you think his points make sense, his job title ("head of strategic futures @openai") means they should all be taken with a massive grain of salt.
That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
I have sonnet do the thinking, deepseek does all the tasks. I've massively reduced costs with this approach.
This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.
And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.
The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story
Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.
It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.
Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)
I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.
Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.
But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].
If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.
But that's a big if we just don't know for sure.
1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...
Or any US hyperscaler with GPUs to spare can decide to serve the models for a reasonable cost/token.
You don't have to send China your data.
I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.
A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
Sure as an indie hacker, you could go download the weights for a Chinese model with a VPN, and then attempt to run it at home by building your own GPU cluster but these large models require quite expensive hardware to run on and so it makes it less likely than anyone would invest that much capital to do something that is illegal. There's no way for them to sell a legal service using those tokens. So it can only be strictly for personal use (the Govt won't care because very few people will have that kind of money and risk appetite). The other option will be that there will be some shady third-party providers in foreign countries who are willing to sell tokens from these models to US consumers knowingly.
So under a ban rest-of-world gets to use cheap open-weight models but American companies/individuals must only use only ‘approved models from US for-profits’? Doesn’t seem like that kind of protectionism will be popular or politically tenable. Not so long ago US chose cheap TVs over maintaining the country’s manufacturing base.
(Despite what you wrote it’s also really hard to imagine that enforcement wouldn’t leak like a sieve. Unser sufficient economic incentives [which are the predicate for the ban], loopholes will be found.)
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.
The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.
IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.
As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.
Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.
I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.
So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.
Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.
You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
An astute Chinese analyst could reasonably forecast that they had little chance of controlling the AI market due to sovereign trust issues, but would also note that AIs are just software.
When the dust settles the US still won't have factories, and the real value of AI models is still going to be embodying them and getting them to do real, consumer facing work.
Perhaps the most striking thing about the AI boom is how quickly the US abandoned the veneer of local manufacturing in favor of more expensive buildings producing nothing you couldn't make anywhere else on the planet...from imported parts.
how is running servers supposed to be 0 cost, while running ai inferrence isn't?
For a SaaS business, running servers isn't free. But compared to the cost of running GPUs for inference that you are selling, it almost is. The company I work for is a SaaS company. We have a single production server. A couple of QA servers. All hosted on Hetzner. Monthly cost for servers is less than $400. This generates a few million dollars a year in revenue.
If we were in the business of selling inference, our cost of providing the service, for the same amount of revenue would significantly higher.
Even large businesses like Microsoft, Meta, Google have operated with similar margins. Cost of running servers, compared to revenue was very low. But inference changed that, in a dramatic way.
A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.
lmao
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is. but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency. and they even plan to build their own inference hardware, too. and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
and you say there is nothing to afraid?
China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.
Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.
So I believe, China will end up owing AI.
I think you hit the nail on the head - right there!
Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
I think it is the right move to protect American interests
Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.
Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.
And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."
[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".
[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
This from OpenAi's Head of Strategic Futures "Some observations on Kimi: It's a very good model! I don't think its performance can be explained away by distillation or anything like that"
https://x.com/deanwball/status/2078133895766114412
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
To win on the AI front by any means necessary.
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
Why wouldn't it be? China is pumping out AI research and researchers at a staggering pace and there is no inherent reason why western models should be better
The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.
> monetizing these strengths
> real advantage over China is freedom
Please tell me if I'm unfairly paraphrasing but these seem to be your main argument and they seem to be oxymorons