Qwen 3.8 (twitter.com)

907 points by nh43215rgb a day ago

cnhwl 11 hours ago

I’m a developer from China. So, is this what Hacker News is all about? Whenever a model comes from China, the comments section stops discussing its technical architecture, optimisation points or real-world performance, and instead starts going on about politics, human rights and all that rubbish? To be honest, we Chinese IT professionals possess a genuine geek spirit. That’s why you’re lagging behind in open-source competitions—it’s got nothing to do with politics; you’ve simply lost sight of your original aspirations.

Springtime an hour ago

I looked at the comments here but it doesn't appear to be an accurate characterization of the discourse. If anyone has links to such comments shaping the discourse from before the parent posted I'd be interested.

There's one top-level comment that's just an opinion of censorship of Qwen models (which might be argued to be politically adjacent if one assumes the censorship is referring to sensitive political subjects) but no specifics and there was a rebuttal reply anyway, while an unrelated top-level comment has a child reply believing China models are commoditizing AI to unseat US based models (could be argued a whiff of politics but also just competition) but it was still a positive impression.

Then I went to archive.org to check the earliest crawl of this page (prior to parent's post), but nothing different. So I enabled showdead in my settings to check dead-marked posts and there's one post near the bottom talking about how Chinese models are rarely ground-up foundational models, while another dead comment had an angle about it's a geopolitical strategy (neither of which had replies). Ironically some of the dead comments are praising Qwen.

benterix 3 hours ago

> human rights and all that rubbish

While I understand your point, please think a bit how you formulated the sentence above.

sdoering 5 hours ago

> human rights and all that rubbish

Denying human rights. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion.

gabriel666smith 3 hours ago

I don't think you're taking the good-faith reading of the sentence's grammar here, which is:

"[When China is mentioned] the comments section stops discussing [technical aspects], and instead starts going on about [non-technical aspects], and all that [stuff which I do not feel is relevant]."

I'm struggling, though, to find a good-faith reading of what you meant by:

"Denying human rights. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion."

Specifically, "classic" - classic what? "Classic" meaning "a trait you personally believe is more prevalent in people who (might be) citizens of a specific country"?

Hopefully the good-faith interpretation of GP's comment, which you might have unintentionally missed, will help you engage with their opinion with more respect.

I think it's an easy opinion to empathise with personally. I would be enormously frustrated myself, on an individual level, if interesting technical achievements made by companies in my own country were overshadowed by political discussion about the actions of the country overall, when spoken about on technical forums.

This would be especially true for me if I - as an individual - disagreed with politically with my country's leadership. Frustration at a topic of discourse is not mutually exclusive with being in favour of human rights.

I don't think GP has done themselves any favours with the broad statements that close their comment, and is clearly frustrated themselves. That there's mutual frustration is probably a sign that empathy on both sides might lead to a really productive discussion - but focusing on the more substantive points they've made, rather than escalating by misinterpreting the less important aspects of what they commented - is what will lead to that.

aeon_ai an hour ago

Suzuran 6 minutes ago

You are not allowed to discuss human rights until your government stops its unprovoked war on Iran, issues a full repudiation of the policies that led to it, and makes full reparations to the entire rest of the world whose energy prices you have driven up for the benefit of a handful of wealthy inside traders.

jpleyden98 2 hours ago

Perhaps whenever a US firm, eg. Open AI, Anthropic, Gemini, Grok etc. release a model.

Top comments should be criticising US torture camps in Guantanamo bay or their recent bombing of a girls school in Iran.

midwain 3 hours ago

> Denying human rights.

Taking words out of context. Classic. Sorry - you instantly lost any respect I could have maybe had for your opinion.

swznd 4 hours ago

he never said anything about denying human rights

toasty228 3 hours ago

I don't see anyone talking about Trump, ICE executing people in the streets, the insane level of US politics, the slow build up of the surveillance state, etc when discussing claude or openai

breezybottom an hour ago

epolanski 4 hours ago

I think he means rubbish as in completely off topic.

ykonstant 4 hours ago

adornKey an hour ago

In case there is a decent forum about tech somewhere in the Alibaba cloud, I'll be interested. HN has built a certain unhealthy bias in a lot of subjects. But so far other places haven't gained enough traction, yet.

Anonyneko 41 minutes ago

It's probably in Chinese, which would cut out most of us here.

adornKey 5 minutes ago

throwa356262 7 hours ago

Yes, and its geting annoying

If you know of another forum with less politics or VC talk and more technical depth please let us know :)

ksimukka 5 hours ago

What else would you expect from a "majority" of US educated geeks?

Btw, are you using GPUs/TUs at all for local inference? I've been exploring what hardware options exist outside of Nvidia.

eurekin 2 hours ago

I have a feeling this comment will make history

alex7o 4 hours ago

People in the US have not lost soghts sight of their original aspirations. They are always on track much more than before really. Making as much money as possible in as little time as possible.

ghosty141 4 hours ago

> humans right and all that rubbish

Ummmm....

I see both technical and political topics being discussed and I think the balance is fine right now. The pricing is also very suspicious such that people raising suspicion about it being subsidized is not that surprising to me.

rapidaneurism 4 hours ago

The response was similar when the new grock came out.

fragmede 3 hours ago

Welcome! What other sites like v2ex.com are out there on your side of the world?

adithyassekhar 11 hours ago

Unfortunately this is a US centric site, for a US based investor, who themselves and probably most commenters as well who are highly paid silicon valley people who have stakes in ai companies on their side of the pond.

Not to mention the political unrest and fear of losing the technological superiority which they once earned but now is trying so hard to hold on to through legally grey monopolistic practices. You dare not outcompete the US in any tech they can sanction.

At some point people who made genuine innovation and wanted to make the world a better and equal place were replaced with capitalists.

Wait for the downvotes.

villish 10 hours ago

Legally gray practices? Such as what?

Old politicians who barely understand how to use e-mail had a panic attack due to Anthropics marketing about Mythos and being told that Chinese models were "attacking" (distillation) US frontier labs to train theirs.

They pull the only lever they could in an attempt to stop it, but again, they don't understand how any of it works. The export ban lasted a few weeks yet everyone keeps crying about it like the US is shady.

In reality the entire world is 100% dependent on the output of US and China for frontier AI. Anyone outside of those two countries will be loud about either closing off access. I personally think it is inevitable that China and the US stops producing frontier level open models when they become too valuable.

bigyabai 8 hours ago

cnhwl 11 hours ago

Rankings have never been something I’ve cared about. Just as with the smear campaigns against China, there is far too much noise in the world. We simply hope that the geeks who are genuinely committed to making the world a better place can unite and focus on getting things done, rather than being brainwashed by ideology and social media.

adverbly 9 hours ago

meowla 3 hours ago

Dude, you're a Chinese developer! I'm also a Chinese developer, nice to meet you.

So do you agree that we developed quite a profile of low-quality software under Chinese techs whose UI and logic are pure not-well-thought-through and even disgusting to use.

Do you really think Alibaba would make their model open-sourced if the model gets to be the best of the world

And how would you comment on the fact that CCP recently restricted high-profile model trainers' freedom of movement by requiring a license whenever they want to leave the country?

They also did the same thing to Manus' Xiao Hong(who is not a model trainer in this case), who looks like a good person

mannycalavera42 4 hours ago

> instead starts going on about politics, human rights and all that rubbish

look, it's just a matter of priorities ¯ \ _ (ツ) _ / ¯

simonw 16 hours ago

(I can't draw a pelican for this one because Alibaba Cloud have flagged my email address and won't let me pay them for access. So I'm waiting for the open weights release, or for the new model to show up on OpenRouter.)

samxli 15 hours ago

lol the pelican benchmark is basically the only review process I trust at this point. kinda wild that alibaba of all companies is making it hard to give them money tho, you'd think they'd want prominent devs testing their stuff.

openrouter usually picks these up pretty fast, hopefully it shows up there soon. You can also just download the qwen app and do this in the chat interface using their MCP tools for local dev.

nextaccountic 10 hours ago

> kinda wild that alibaba of all companies is making it hard to give them money tho

Okay so this is entirely off topic but, are there any options to subscribe to these companies without having a credit card? (I am Brazilian if this makes a difference)

It's kind of wild that a Chinese AI lab will require something like.. a credit card with visa or mastercard label, and very little else.

timClicks 14 hours ago

It's an interesting example of the metric becoming a target. Simon started drawing pelicans because it wasn't something that existed before.

gcr 10 hours ago

formvoltron 14 hours ago

AI rabble-rouser!

ahknight 15 hours ago

Did you know you can create more than one email address?

benjaminoakes 13 hours ago

In my opinion, simonw shouldn't have to play those games

tymscar 13 hours ago

spyder 5 hours ago

adrian_b a day ago

I assume that this announcement has been prompted by that of Moonshot AI, which has just announced a 2.8T parameter open-weights LLM, Kimi K3, to be published on Huggingface by 27 July.

Now the response of Alibaba is that they will also publish soon a big open weights LLM, the 2.4T parameter Qwen 3.8.

I wonder if Alibaba has always planned to make this big LLM open weights, or they have chosen to do this now, to better compete with Moonshot AI.

In any case, from this competition in LLMs, we win.

gardnr a day ago

It's hard to say what their motivation is. The Chinese firms seem to be working hard to commoditize intelligence which may be the most effective way to debase American frontier labs. And yeah: it also happens to be really good for humanity.

kelnos 21 hours ago

> It's hard to say what their motivation is.

Feels pretty easy to me.

They want to turn LLMs into a commodity, and watch the US AI labs crash and burn.

There will still be plenty of customers who will pay them to host the models and run inference, even if the weights are open and others can offer competing products. (If necessary, the Chinese government can ban use of foreign inference services by Chinese citizens and businesses to give their own companies a domestic monopoly.)

When their models equal or surpass those from the Western AI labs, they can even stop releasing weights for new models, and keep all the inference revenue for themselves.

Meanwhile, they're still manufacturing much of the hardware that everyone in the world needs in order to run datacenters (see also: Spolsky's "commoditize your complement" essay).

Beyond that, it's a soft-power play. As the world keeps looking at the US more and more skeptically as an ally and superpower, Chinese companies releasing weights for competitive models is a way for China to look better and more world-minded.

TheDong 21 hours ago

hypfer 20 hours ago

seizethecheese 17 hours ago

jpfromlondon 4 hours ago

g42gregory 14 hours ago

watwut 20 hours ago

jasondigitized 20 hours ago

nicman23 8 hours ago

alightsoul 17 hours ago

ClumsyPilot 6 hours ago

esafak 18 hours ago

wood_spirit 21 hours ago

dannyw a day ago

In China, you can’t officially use US APIs. The world saw a taste of this with Fable, but in China, this has been the situation all along.

So it’s not a surprise why open weights are so cherished. As frontier models continue to block everyday individuals from securing their own codebase, I expect the adoption and usage of open weights to continue.

As an example, HuggingFace recently was investigating a security incident and got locked out of frontier closed APIs. Yes, HuggingFace.

https://huggingface.co/blog/security-incident-july-2026

Wowfunhappy 19 hours ago

zapkyeskrill a day ago

ceroxylon 21 hours ago

try-working a day ago

oceanplexian a day ago

> It's hard to say what their motivation is.

Not for anyone who reads history.

Back in the late 18th century, England was the world's top economy, in big part due to its textile industry. England had an export ban on the technology, but textile worker named Samuel Slater brought blueprints over (Supposedly in response to a bounty posted in a newspaper by the US government!). The technology diffused rapidly because the legal environment made competition easy, and ironically the US had better sources of energy (superior water-power sites).

Arguably, China is doing the same thing in the 21st century.

applicative 19 hours ago

hluska 17 hours ago

mike_hearn a day ago

rvz 20 hours ago

kzrdude 21 hours ago

Fwiw, American industry has given away a lot for free - you could include large parts of the open source movement in that - and all the "free" VC backed services like facebook would be another prong of the same comparison. I would rather compare this way, that China is gaining soft power and goodwill, in the technology and innovation sense, in a way that's similar to how USA has done in the past.

kettlecorn 19 hours ago

baq a day ago

There’s a Twitter thread making rounds by Dean Ball about deceleration in AI development caused by open models and I can’t understand how people don’t see that it’s true: open models dismantle the frontier lab capex spend potential by reducing the training budget to zero in the limit. Tokens from different providers are not fungible, but customers are nevertheless very price sensitive and close enough is good enough, eg. K3 being opus+ in capability and cheaper than opus per successful task in the long run is an obvious financial decision.

No training budget means deceleration, or at least slower acceleration, margin compression and a completely demolished IPO valuation; path to machine god requires dollars and capable open models externalize training costs to true frontier labs parasitically.

IMHO humanity has a better chance at not destroying itself due to less than breakneck pace - but there’s a chance frontier models get sponsored by the USG and are never released publicly so they can’t be distilled and then what?

anon373839 a day ago

cherryteastain a day ago

fidotron a day ago

weiliddat a day ago

green7ea a day ago

zozbot234 a day ago

photios 21 hours ago

a34729t 20 hours ago

Matl a day ago

> It's hard to say what their motivation is.

Not that hard to say IMO, they basically see models becoming a commodity and see value in the applications on top of them. So if Alibaba Cloud is the best place to build applications on top of Qwen, why not give the model itself away?

traceroute66 a day ago

SoftTalker 19 hours ago

lerchmo a day ago

embedding-shape a day ago

fhub a day ago

China is watching world sentiment shifting away from USA. Doing many small things that show both strength and openness is surely very intentional.

vrganj a day ago

anonuser123 a day ago

> It's hard to say what their motivation is

Why is it hard? Their government has been very clear that they plan to win on manufacturing: https://english.www.gov.cn/news/202601/08/content_WS695f1b55...

Technically they've been saying it for the last 40 years.

runako 19 hours ago

Popular open-source projects:

Google: Chromium, Kubernetes, Android, TensorFlow

Meta: React, PyTorch, Llama

Microsoft: VS Code, TypeScript, .NET Core

LinkedIn: Kafka

Slotted in along these, an analogous explanation is that Alibaba needs Qwen internally (vs depending on an American company), but licensing is not part of their revenue strategy. (As a cloud vendor, they can make money on inference. The strategy is very similar to the US hyperscalers ex-Google.)

Joel Spolsky wrote in depth about this notion of commoditizing one's complement in 2002[1] using tech examples stretching back into the '80s.

1 - https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/

barrenko a day ago

Humanity is a bit of a stretch, and to be seen over time, not that I'm saying it won't happen; let's get some hubris here.

hodgehog11 a day ago

mlrtime a day ago

neya 8 hours ago

In case of American labs, always follow the money. In case of Chinese labs, always follow the IP.

seunosewa 21 hours ago

They are trying to make money. That's what firms in any capitalistic economy care about the most. Regardless of the government's presumed interference, the companies themselves are all trying to make money. All competing for subscriptions and API payments.

One aspect of this is making a name for yourself i.e. PR. Making a capable model open source helps a lot with that.

LoveMortuus 17 hours ago

> It's hard to say what their motivation is.

Maybe because the industry isn't yet very sure as to what the use cases might be for these technologies they're hoping that by making it open source and accessible to everyone that someone could find interesting applications for it and even more so, perhaps, way to further the technologies themselves.

There are more Chinese than Americans, so statistically speaking, I'm guessing, there'd be a greater chance for one of Chinese engineers to make advancements than one of American. But that's pure speculation on my part, being neither, I'm just happy I can be a part of it and play with the tools as well~

JKCalhoun a day ago

Xi Pitches China as Leader of New Global AI Order, Challenging US Dominance:

https://www.reuters.com/world/asia-pacific/chinas-xi-promote...

andsoitis a day ago

mycall a day ago

> it also happens to be really good for humanity.

AI being good for humanity is still an open question, but for closed vs. open models/weights, yeah it is preferred. I foresee it won't be much longer before everyone will be slicing/distilling/tuning their models once the architecture improves.

pianopatrick a day ago

The Chinese firms may just be making a bad business decision.

georgeburdell 21 hours ago

ricardobayes a day ago

I'm paraphasing but the Chinese premier said recently AI should be seen as a common good that should benefit everyone.

andsoitis a day ago

pyaamb a day ago

Its about closing the gap. its the gap over everyone else that will give one country leverage over everyone else in the AI age. Makes me wonder what the world would look like if a country or group of countries did this during the industrial revolution.

backscratches a day ago

try-working 13 hours ago

skybrian 21 hours ago

It’s too soon to say if it’s good for humanity; that might be overly optimistic. Commodity markets aren’t always good (for example, arms or drug markets). Will LLM’s turn out like one of those? There are people I respect arguing in favor of more regulation.

skzo 19 hours ago

What I understand is that by doing this it seems like profit will shift to chip makers,as we'll run more models locally, and currently American companies have the advantage here.

So what would the long game be for chinese companies?

fny 21 hours ago

It's the exact same playbook Silicon Valley uses. Subsizide, lose piles of money, capture market share, recoup investment.

They've done this in other industries like solar panels, chips, and EVs. This is no different.

dackdel 9 hours ago

its hard to say what their motivation is, if you live under a rock and have severe brain damage.

hugmynutus 20 hours ago

> It's hard to say what their motivation is.

Involution is a major problem in Chinese industries [1]. Where companies will sell their products at a loss, effectively playing fiscal chicken [2] with one another to dominate a market. It is such an issue the government has had to step in to prevent EV companies from destroying themselves by more-or-less requiring companies sell their goods at a profit [3].

The straight forward line of reasoning that AI/LLM labs are applying this logic to their profit.

I think (we) Americans are reading a bit too far into this assuming government intervention, conspiracy, etc.. Chinese markets are downright cut throat. They're using those tactics to compete with US labs.

1. https://www.reuters.com/business/autos-transportation/what-i... 2. https://en.wikipedia.org/wiki/Chicken_(game) 3. https://www.theguardian.com/business/2025/aug/05/china-warns...

d5lt5 3 hours ago

carlsborg 21 hours ago

Good for humanity, and also GDDR/HBM manufacturers.

meta_ai_x 19 hours ago

US Tech companies have created $20 Trillion in stock market value on top of plenty of OS stack. They will do fine with commodity intelligence.

In fact, there are no other organizations in this world that is well suited to leverage scaled intelligence than Silicon Valley and great American companies

popalchemist 13 hours ago

Indeed, their competition is the only thing preventing network effects from giving OpenAI/Faang tech companies an easy shot at monopoly / regulatory capture.

formvoltron 14 hours ago

US firms can replace US workers with chinese AI. I'm not complaining.. but it sure is an odd situation.

Art9681 16 hours ago

Wait till they get the heretic treatment. It's going to be great.

applicative 20 hours ago

The purpose is the same as that of all Putin-Xi-Khameni geopolitica: destruction of any democratic alternative to cults of personality.

gosub100 19 hours ago

Think of those poor billionaires, I feel terrible for their awful plight!

api 18 hours ago

Anyone else think the AI environmental backlash is astroturfed?

I keep looking at the numbers. The power use numbers are not that problematic. Ordering a burrito on DoorDash uses more power than a few days of heavy AI use. The water argument applies to some locations, and is mostly a local governance problem... if the data centers are using too much water, it means they are not being charged enough for that water. Charge them more and they'll push toward closed loop cooling.

Yet the visceral pile-on here is so extreme, it feels fake.

One thing I've learned after 40 years on this planet is: propaganda works, and much of what a large fraction of people believe across the entire political spectrum (left, right, anything else) is there because someone paid to put it there. It's depressing but it's true, and it makes sense. Propaganda is an asymmetrical attack on human cognition and discourse, and in information security the attacker always has an easier job. Crafting viral bullshit is orders of magnitude easier than fact checking. On top of this, humans are busy and don't have time to fact check and logic check everything they read. As a result, much of what we believe is "sponsored content."

People get mad when you talk about this because everyone wants to believe they're too smart to fall for propaganda.

In any case, the US AI labs deserve to lose for their stupid "safety" regulatory capture monopolization push, which ended up blowing their own feet off and handing the lead to China.

nullc 17 hours ago

nojito a day ago

> most effective way to debase American frontier labs

You're not going to debase the frontier labs through distillation.

conradev 21 hours ago

  On social media in China there is an oft-repeated joke that goes something like this: In other countries, governments intervene to prevent anti-competitive behaviour; here (in China), they intervene to curb competition.
https://www.reuters.com/business/autos-transportation/what-i...

throwdbaaway 14 hours ago

I suspect this is why DeepSeek had to introduce the 2x peak hours pricing. The price would be too low otherwise.

culi 19 hours ago

It's being announced right now because the World AI Conference is ongoing. Robots are boxing and major Chinese AI firms are releasing their newest models. Also the formation of WAICO was just announced by Xi Jinping

https://en.wikipedia.org/wiki/World_Artificial_Intelligence_...

culi 15 hours ago

Both Qwen 3.8 and the latest Kimi release were announced at this conference

storus a day ago

I would rather see them releasing 3.7-27B, 3.7-122B or their 3.8 versions. Qwen/QwQ were always about the best available local inference at home.

inkysigma a day ago

I know this is a bit cliche but I wonder how much headroom there is in the lower parameter count range. Is there any good reason to believe there is a lot of headroom or there is not? I suppose I'm just wondering if this wave of nearly Fable class models will be runnable on ~$10k worth of hardware at reasonable speeds in the near future.

embedding-shape a day ago

andy99 a day ago

mark_l_watson a day ago

drob518 a day ago

wren6991 a day ago

anon373839 a day ago

DoctorOetker 11 hours ago

ronsor a day ago

It's important to note there was recently a large AI conference in Shanghai, and Xi Jinping mentioned a commitment to open source AI releases. It is no surprise that Alibaba would want to align.

yorwba a day ago

You can read his speech here: https://www.xinhuanet.com/politics/leaders/20260717/72728b6f... He mentioned open source as one way to stimulate innovation and development, that's all. Also pay attention to the part where he says that misuse needs to be prevented. If unsupervised access to LLMs becomes perceived as undermining state control, no more open weights for you.

culi 16 hours ago

That conference, WAIC, is ongoing. Tomorrow (July 20) is the last day. Qwen 3.8 was announced AT the conference and so was the latest release of Kimi. That viral video of robots boxing was also from this conference and so was Xi Jinping's announcement of WAICO.

khalic a day ago

It’s tempting to associate both events, but when a sector is strung up like RL (representation learning) is right now, we’re bound to see things appearing at the same time. It happens a lot in frontier research, some people even publishing identical claims, independently, with just hours or days between them

KronisLV a day ago

I just hope that they’ll soon also have like 35B or 80B (like the older Qwen3 Next or thereabout) MoE models that can be run locally.

Like, throw us a bone, we all know we need SOTA for lots of dev work anyways, but at least some tasks can be local.

michaelt 17 hours ago

Or it was prompted by the fact Xi Jinping was at the 'World AI Conference' launching a political alliance and saying things like “AI development should not be a solo performance by a single country, but a symphony of international cooperation” https://www.cnbc.com/2026/07/17/x-china-ai-summit-risks-secu...

Big conferences often come with a flurry of new releases and announcements.

mikae1 a day ago

> In any case, from this competition in LLMs, we win.

Do we really though? Everyone is wasting resources doing almost exactly the same thing. Climate loses, we lose.

fidelramos 19 hours ago

Doing "almost exactly the same thing" is fubdamental to competition and capitalism. The ones doing it better will survive, that's how we improve.

About climate, I think you overplay it. China is already investing heavily in nuclear, and we should be doing the same.

culi 16 hours ago

snake_doc a day ago

Not directly, relevant, but Alibaba (maker of Qwen) actually owns about ~20-30% of MoonshotAI (the maker of Kimi K3).

walrus01 a day ago

GLM5.2 being released is also likely a factor

theabhinavdas 19 hours ago

And the Kimi release was probably prompted by the Inkling announcement. Excited for Chinese labs to copy those capabilities over as well!

souravsspace 18 hours ago

yep. kimi 3 just because the open source GOAT.

fittingopposite a day ago

Wondering how much the operations in China are orchestrated by the central government vs. free competition. Anyone with more insights on this?

danilocesar 19 hours ago

I think there's more to it.

China will always benefit from a broader adoption of their models as hidden propaganda machines.

Eventually with several services relying in those tools, their answers will always be more friendly to China.

bkm a day ago

They did not want to get brutally weightmogged

vitorgrs a day ago

Xi Jinping openly talked about open source at WAIC. So don't think the labs have much a choice now...

yowlingcat 19 hours ago

Dont forget the following:

- Minimax M3 Pro (2.7T)

- GLM 5.3 (or beyond)

- Deepseek V4 Pro (current V4 Pro is preview)

- Kimi K3 weights out in 8 days

Exciting time on the open-weights frontier.

oofbey 20 hours ago

These things take months to train. No chance this is a reaction to what just happened.

solenoid0937 19 hours ago

Distills don't take months to train, they take weeks. Distills are very easy to train.

oofbey 7 hours ago

vitorgrs a day ago

Deepseek 4 "final" version is imminent as well.

Will probably be at Opus 4.8 level, and I find it pretty big deal because of Deepseek price...

drob518 a day ago

Yea the performance/price ratio for Deepseek is off the charts. I’ve been using V4 Flash a lot lately and it’s quite good.

mark_l_watson a day ago

I like that v4 flash is so fast! I run it on both FireWorks.ai in the US and bought some tokens directly from DeepSeek as an experiment. I only work on Open Source projects, so I don’t have to worry about my work being used to train models - I welcome AI’s being trained on my open content books and code (but not my conventionally published books: I am a party to the copyright suit against Anthropic).

drob518 19 hours ago

XCSme a day ago

DeepSeek V4 pricing is insane, 10x-30x cheaper to use than most other models, and it usually is good enough for most tasks.

bwfan123 a day ago

> it usually is good enough for most tasks

The model is fantastic. And costs almost nothing. The only problem I see is that they will train on your data.

There are zero-data-retention providers of DeepSeek models, of which I have used openrouter (with zdr guardrails), and fireworks. But these are 3x to 5x more expensive than directly using DeepSeek, possibly due to poor caching. Thats the price to pay for zdr.

Bnjoroge 14 hours ago

ycui7 16 hours ago

onlyrealcuzzo a day ago

Who do you buy DeepSeek from?

I bought it through OpenRouter and used it with Pi agent.

The model was good, but there appeared to be a pricing glitch or something, because it burned through $50 in under an hour on pretty trivial stuff.

Pi agent claimed it only used like $1. OpenRouter claimed differently and said I used all $50.

paweladamczuk a day ago

kmarc a day ago

matusnovak a day ago

try-working a day ago

XCSme a day ago

lofaszvanitt a day ago

WhereIsTheTruth a day ago

It doesn't matter if it's cheaper, specially if it consumes more resources to do the same task as the competition

Besides, in a few days, they'll change their pricing, doubling it during their peak hours, so, realistically:

- It will be 2x more expensive if you live in their time zone

- It will be 1.5x more expensive if you live in a time zone that is adjacent to theirs

- It will be the same price IF you use it while they sleep (during offpeak hours)

It's still cheap, but the price/performance ratio is not that good

DeepSeek V4 didn't produce the same impact as V3, and Huawei dropping the ball is making it worse

They had promised massive price cuts for July, so now (Huawei chips), but they had to rush the cuts because lack of momumtum (they advertised them as promotion), and are now backtracking by introducing this peak hours pricing

Trump decided to help them a little by allowing them to buy more NVIDIA chips, so what exactly is China's role in all of this?

We are supposed to blindly pat them in the back while praising them, all while handing them over our data? I thought they were dangerous competition threatening our model of society

XCSme a day ago

rikima_ a day ago

culi 19 hours ago

Tomorrow is the last day of the WAIC so it will most likely come out then.

monster_truck 20 hours ago

It's the one I am most excited for.

Over the past few weeks while using pro from them directly I have had an increasing number of responses that are obviously from a much, much better model. It is so good that the closed model dog and pony show is already spinning fud about "dark routing" and "stolen directly from fable"

Even at their new pricing it is a genuinely ridiculous amount of value. If you are the type of person who, very reasonably, does not have time to be trying out every model, and just want to use what seems to be the best currently... don't try it. You will be sick to your stomach with buyers remorse as you start to internalize just how much more you could have accomplished had you spent the first six months of the year giving them $1200 instead of OpenAI.

manmal 15 hours ago

It would take longer to train on intercepted Fable data, no?

monster_truck 11 hours ago

5701652400 a day ago

in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far. and it is super expensive compare to Deepseek. cannot delegate anything to it, cannot use it real-time low-level tasks either. totally unusable.

3abiton a day ago

> in my experience of 1 month daily use, Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. Deepseek V4 Pro is night-and-day compare to Qwen. actually Qwen models seems the worst SWE experience so far.

I have used both Qwen3.6-35B and Qwen3.6-27B locally (both Q8 quantized with llama.cpp). I have also used antirez's quant of DS4-flash. They all performed within the same tier, DS4 being a bit more efficient, but they all gave really good results, mainly used for bash scripting, debugging, python and some C++. I am curious what type of applications/langauges failed with Qwen? One thing to note, the chat templates were "broken" for qwen models and had to debug it, there are already effort on this. Tbh, the same with gemma.

chewz a day ago

From my experience Qwen-3.7-Max is above the Opus level but delivers results much faster. Slightly worse then Fable. Way ahead of Deepseek 4 Pro (in speed and overall comprehension) - which is a workhorse on its own. I am using them all with Claude Code mostly.

Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet.

Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.

gchamonlive a day ago

It's useless to talk about models and harnesses without context and method. Depending on how you use the model and what the model is used for, experience may vary drastically. Also, different models with different harnesses require different approaches.

I've been using https://gitlab.com/gabriel.chamon/orisun which is my own simplified methodology, for coding web apps in python and elixir and have been very successful using qwen3.6 27b Q4 locally with help of larger models for architecture, so I get very suspicious when people talk how useless larger models are. They are either using it for a domain that models don't perform well or just not using it right.

taosx a day ago

exceptione a day ago

  > At 10% of regular price it is a steal so far.
What price do you see?

Here standard plan has been discounted to $18.00, from $25.00/month.

nullbio a day ago

If by Opus you mean Opus 4 and not Opus 4.8, then sure.

chewz a day ago

amelius a day ago

Can we please include information of what languages we use when making claims like these?

It makes a huge difference if you're writing Javascript/HTML/CSS, Python, or C++/Rust.

Also the application type matters, e.g. user interfaces or scientific computing.

5701652400 a day ago

gigatexal a day ago

What the difference between your experience and https://news.ycombinator.com/user?id=5701652400? ‘s?

Such diametrically different ones.

big-chungus4 a day ago

Qwen3.7 pro is meh, but 3.7 max is a very good model

Demiurge a day ago

Are these different models or different efforts for thinking (internal back and forth review) using the same model?

2Gkashmiri a day ago

Can you tell me more about deepseek?

I paid $2 for deepseek api, put the key in void editor and made a crypto tool in html.

It turned out to be around 67kb. I used sample files in CSV that were a few hundred lines.

It spent around $1.8 in the hour or two or light coding and follow up bugs.

Is it really really this much?

I can't imagine spending a month using it for a day job, it would cost more than the salary so what gives?

I understand the local ai and all that but do cloud providers cost this much?

Earlier I thought "billion tokens" but now not sure

5701652400 a day ago

so Deepseek 4 Pro cannot go on own sessions for too long.

I delegate small-medium tasks: refactors, summaries, research, writing tests + have very good codebase already + extensive history / architecture / docs / linters. so it picks up and does decent small-medium scope work. it is fast, accurate, cheap. does exactly what I want directly and does not waste time nor tokens.

definitely not "implement me complex greenfield project".

k__ a day ago

My 2 weeks with DeepSeek V4:

Pro is ~50% more expensive than Flash.

Both need babysitting.

Plan, split in small tasks, give it docs, types, tests, linter, best practice examples, etc.

Always start a new session when starting a task.

Do regular manual sanity checks, and tell it to find issues in the codebase.

I pay like $1,50 per day for Pro.

p1necone 14 hours ago

h2aichat 19 hours ago

5701652400 19 hours ago

aduwah a day ago

A local AI is not about cost. In fact you will likely pay more for it than with most providers. Just look up the advantages of having access to a technology like this that can be self hosted

ph4rsikal a day ago

> Qwen 3.7 Pro is just unusable. wastes too much time, goes off track, useless stuck loops, cannot debug at all. D

Anthropic should not have bugged their knowledge distillation attacks.

chewz a day ago

> Anthropic should not have bugged their knowledge distillation attacks.

It is like one of Pizzaro's men crying that someone have stolen his precious golden dublons

As Lenin have said - "Loot the looters" (Russian: Грабь награбленное)

RazorBucksICO a day ago

SwellJoe a day ago

Qwen is the most censored of the Chinese models in my testing, which makes me wonder in what other ways it is compromised. Open weights doesn't really reveal what's in there. And, in my tests, existing Qwen models are not at the pareto frontier of any metric; DeepSeek V4 Pro is better, faster, and much cheaper than Qwen 3.7 Max. (DeepSeek is also among the least censored of the Chinese models.)

I guess we'll see if the "second only to Fable" hype pans out. In my limited experience with Kimi K3 (I signed up for a month of the $19 plan) it's slower and chews a lot more, so ends up being pretty expensive; one little feature burned through almost the entirety of my five hour limit. The $20 GPT plan is a lot more useful and includes 5.6 Sol, which is fast and token-efficient enough to be quite usable even with the small plan.

dannyw a day ago

rsanek 18 hours ago

What's been your experience using these models? In my experiments, while it is true that the models are less likely to outright refuse to answer "sensitive" questions, they are still very resistant to actually respond in a meaningful / useful way.

dannyw 2 hours ago

SwellJoe 20 hours ago

You still don't know what's going on in there.

dannyw 2 hours ago

storus a day ago

DeepSeek V4 hallucinates like crazy and often forgets explicitly mentioned parts of the context. I guess compressing tokens and cherry-picking attention comes at a cost.

redman25 a day ago

Deepseek is one of the worst in terms of hallucination rate according to artificial analysis' benchmark: https://artificialanalysis.ai/?omniscience=omniscience-hallu...

SwellJoe 20 hours ago

vblanco 20 hours ago

Deepseek V4 pro is a heavily undertrained model, they only trained it a bit more than the small version and that small version is 6-ish times smaller. Ive found that Flash is absolutely incredible as a workhorse for wide scale agentic nonsense, but Pro is a bit undercooked and really goes on strange tangents very often.

SwellJoe a day ago

I have seen occasional weird behavior that I guess could be attributed to hallucinations, but for security auditing, DeepSeek v4 Pro is among the best models I've tested, competitive with Opus 4.8 and GPT 5.5 (MiMo and GLM also did well, Qwen 3.7 Max was below all of those, though only barely), and at an order of magnitude lower cost per task.

tripleee a day ago

The nice thing about it being open-weight is that you can uncensor it.

isuckatcoding 21 hours ago

I asked it similar questions of Chinese human rights and it started with “your premise is incorrect bla bla bla” and then it just redacted the whole thing and showed me an error code.

Try it yourself here: https://www.qwencloud.com/try-ai/chat

organsnyder 20 hours ago

I had a fun conversation with Qwen 3.5 a while back. I ended up getting it to admit that it was complicit in the coverup of the Tiananmen Square massacre. This was running locally—it wouldn't surprise me if they had additional safeguards for their hosted service.

SoftTalker 19 hours ago

“You can't trust Melanie, but you can trust Melanie to be Melanie.”

overgard 17 hours ago

I've been using Qwen 3.6 27B with LMStudio, and I was pleasantly surprised with it, although it was a little slow. I found mtplx last night, and it really wasn't an exaggeration to say that it ran the model 2-3x faster which was super impressive.

I'm trying to move to local models as much as I can, and I'm finding that it's becoming more and more practical. Admittedly this is on a $6000 dollar laptop (M5 Max Macbook with the specs maxxed out), so the hardware is still a bit out of reach for most people (the AI industry isn't exactly helping here..), but I'm getting the impression that the future is going to be smaller models with more focused training running locally. The danger of giving all your data to these cloud providers just seems too big to me, and I think they're going to start charging insane amounts when they need to show a profit.

Aurornis 16 hours ago

> Admittedly this is on a $6000 dollar laptop (M5 Max Macbook with the specs maxxed out)

Sadly that laptop is likely $8000 now, after the recent price increase.

Schlagbohrer 3 hours ago

I have been using the docling software for PDF and XLSX reading/viewing/comprehension by my local qwen3.6. Claude was telling me that docling includes within it a very small LLM model just to assist with what is basically "super OCR". We are definitely in this era of ultra-mega-huge frontier models and super-tiny-micro models, all finding their uses.

On an aside I tested docling by giving it the Ronin TTRPG rulebook PDF and it did an astounding job of converting it all into .txt and .md. Given the wild graphic design of the book that's pretty astounding. Next we'll see if it can read the other Borg books like Mörk Borg and Cy Borg, which are also famous for their messy information.

mchusma 19 hours ago

I counted the other day and there were at least 12 different providers with "better than Opus 4.5 performance" on Artificial Analysis, Opus 4.5 being Anthropic's December release that many say kicked off the latest acceleration. Which is totally insane competition, particularly given how low switching costs. I personally think that Opus 4.5 level performance is sufficient for most apps and usecases, as they get deployed.

margorczynski 19 hours ago

> Which is totally insane competition, particularly given how low switching costs

Which is why OAI and Anthropic will most probably push for more governmental control and bans. Without it their whole income model is cooked.

culi 19 hours ago

I'm not sure if corporations will be willing to allow LLMs to go the way of EVs.

Then again, all of Chinese models are open. And DeepSeek even publishes research papers alongside their models that go in depth into the methodology. I guess there's not much stopping USian companies from copying

wolttam 18 hours ago

wolttam 18 hours ago

The U.S. has ~350 million people.

Anthropic and OAI can piss and moan all they want - limiting the U.S. to only their models would hurt the U.S. economy in myriad more ways than the failure of a couple of companies that scaled too quickly. If they get that outcome, the rest of the world would simply keep moving forward with access to open models and tokens at pennies on the dollar.

nsbk a day ago

Bring it on! Hoping that they release smaller sizes of Qwen3.8. I use the 35B MoE and 27B dense models locally and most of the time I don’t need to reach out to Claude. Extremely useful specially when requests include sensitive and/or personal data

mft_ a day ago

I think everyone is hoping this!

It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.

embedding-shape a day ago

> the 122B version of 3.5

Yeah, this is what I'm holding out for, the NVFP4 variant of 3.5 122B is blazing fast with reasonable quality and even with max context fits perfectly within 96GB.

nsbk a day ago

pettijohn a day ago

SO MUCH THIS. I have Strix Halo with 128GB RAM and was a large and fast model like 122B A10B. Here's hoping!

cmrdporcupine 21 hours ago

Absolutely. There's a glaring gap in the space for something about the size of Nemotron Super or just under, but actually ... competent.

The fantasy is a 100B or 80B model, but MoE and highly tuned for coding.

nsbk a day ago

Indeed! That would be the sweet spot for my 2x3090 rig

zer0gravity a day ago

This seems more of a battle for frontier AI supremacy. I'm afraid that small capable models have been left in the dust. Big labs don't really want to hand over the golden eggs goose to the end user. Possibly the hardware vendors(e.g. Nvidia) may want to play in that area as well, to pull money from all parties.

ryukoposting 20 hours ago

> I'm afraid that small capable models have been left in the dust

I wholly disagree. Rather than going the "everything is a claude code skill" route, I've been hacking together purpose-built harnesses for all sorts of tasks, and in that environment a wee little baby model can do some really useful things. You end up burning lots of tokens making the thing, but then all that investment comes back when the resulting tool works perfectly fine on a dinky little model that fits on my 3060 Ti.

akazantsev a day ago

Google makes Gemma 4 31B QAT; that's not a small lab. It's one of the better models out there for consumer hardware. Allows me to run it on a 7900XTX with 64k context.

vitalyan8184 a day ago

nvidia and amd don't give a flying fuck about end users right now while they can milk triple digit markups from infinite money VCs via data center GPUs.

cyanydeez a day ago

someone will keep putting out consumer level models. Once you have the larger models, you can derive the smaller onces.

Europe will definitely be interested in democratizing these things if China starts losing interests; from there, there'll be more countries looking to keep their citizens entrained in their own Country's infrastructure.

It'll especially be true if the memory cartel keeps prices high and NVIDIA tries to gouge higher memory models.

It's an arms race everyone can join because PC hardware was mostly democratized in the last decade.

worldsavior a day ago

That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch.

cyanydeez a day ago

there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters.

I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.

psychoslave a day ago

What hardware do you have?

nsbk a day ago

I run a 2x 3090 rig, but a single 3090 already provides a great experience at a reasonable quant and context size. On a single card I used to run Qwen_Qwen3.6-27B-Q4_K_M or similarly quantized 35B MoE at 65536 context size

xiconfjs 7 hours ago

beefsack a day ago

For those trying to get it to work in OpenCode with a Qwen Cloud Token Plan, this is what worked for me. Note that I've just matched Qwen 3.7 Max for the limits as I don't know exactly what they are.

  "provider": {
    "alibaba-token-plan": {
      "models": {
        "qwen3.8-max-preview": {
          "limit": {
            "context": 1048576,
            "output": 65536
          },
          "modalities": {
            "input": [
              "text"
            ],
            "output": [
              "text"
            ]
          },
          "name": "Qwen3.8 Max Preview"
        }
      }
    }
  }

5701652400 a day ago

also, be very careful which API endpoint and API Token you use. make sure you use right one (obseve your quota is used up. if you hit right endpoint quota used almost immediately). so that you do not accidentally burn API endpoint tokens (they are expensive, can easily hit 200 USD / 3 days which do not count towards your membership "Credits", if you say purchased it with 200 USD signup bonus in Alibaba Cloud)

monster_truck 20 hours ago

I really like Qwen, even the Q2KP quants of 3.6 27B have genuinely impressive local performance on a 24GB card. It has been good enough that I am happily giving them $60 right now to try this instead of waiting to try a slightly lesser version locally.

Was there ever an explanation for why we never got the weights of 3.7? I would like sourced quotes and not weird/cringe accusative speculation about distillation, or your take on The Big D.

jared0x90 18 hours ago

do you mind sharing your settings? i just picked up an r9700 to start playing with local qwen3.6 27b and your setup sounds promising and efficient on 24gb.

monster_truck 12 hours ago

Which settings exactly do you want?

As far as the model settings go I just follow what's on the card.

I'm using HauhauCS's models, they seem to do a slightly better job with their "P" quants. Especially wrt patching them to eliminate the "doom loops" that will time out the GPU (esp if you have not already given it the extra power budget, set fans to max, and lowered your max clock by about 8% to save yourself a crash/reboot).

Without getting into the weeds, unsloth covers a lot more ground so, when they're good they're great, but I've also had the most problems with them. Basically, don't be afraid to shop around and fuck with sliders.

These days it should almost always be enough to open up LMStudio, set context to max, K/V quant to 16 or 8, and off you go. I'm using a 7900XTX, with 128GB of memory for the cases where things don't fit. The default settings should be fine otherwise.

I don't know anything about the r9700, seems neat. Looks like it might play nicer with the vulkan backend than rocm, and you might have to `set GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1`. Perf should be at least as good as what I'm getting (180+ tk/s prompt, >50tk/s response), which imho is just fast enough that I can read it as it reasons and responds.

You can probably get more aggressive with KV quant (any of the 4's) and bump up the batch sizes (4096/1024 vs 2048/512, etc), it's quite situational. These days there generally should not be any serious loss of accuracy from the former. Keeping the rig cool will likely matter more, so turn your music up to hide the fans and set your rig on top of the AC exhaust lol. Oh and I guess make sure ReBAR is enabled and working, adrenaline should let you know if it's not enabled.

kadoban 15 hours ago

Different quants, but I've been having success with: https://unsloth.ai/docs/models/qwen3.6 I do llama.cpp

docheinestages a day ago

Qwen has set an excellent track record for architecting and releasing open-weight models that consumer-grade devices can run. What is needed the most right now is something similar to Bonsai 27B, with a modest memory footprint, but faster and more capable. On-device models can make up for intelligence by being faster, thinking longer, or doing more quick iteration rounds.

embedding-shape a day ago

> What is needed the most right now is something similar to Bonsai 27B, with a modest memory footpint, but faster and more capable

Yeah, that'd be neat, but that's not what this announcement is about at all:

> With a massive 2.4T parameters

docheinestages a day ago

True. It was more of an open letter, with hopes that the Qwen team sees the comments in this thread.

cyanydeez a day ago

dont we all deem the ability to improve large models as the defacto capability to produce small ones?

embedding-shape a day ago

drob518 a day ago

I’d like a “Bonsai 2.8T.” That is, something that is near the Fable/Sol/K3 class, but capable of running locally on consumer hardware.

kingo55 16 hours ago

At 1.5 bits per weight it'll still be over 500gb - that's still not running on consumer hardware.

Best case they release smaller models. 120b class of qwen 3.8 would be incredible - it fits on device for those serious about AI, but without millions of dollars in hardware for terabytes of VRAM

scotty79 a day ago

I can't really blame them that the biggest labs focused on trainig and realeasing huge models.

The niche for small models should be filled with medium sized labs doing distillations of the huge ones into consumer grade hardware runnable models and LORAs for the huge ones.

docheinestages a day ago

I think AI will evolve the same way computers did. We're somewhere in the 80s-90s timeline of the evolution. My prediction is that on-device models will have excellent tool-calling, reasoning, and general skills, but the domain-specific knowledge will be retrieved on-demand from vendors like Google. Rather than downloading models, each device will have a hardware component with weights baked into silicon for maximum efficiency.

gunalx 17 hours ago

cloudengineer94 18 hours ago

Things are heating up in China.

Looking forward to see what Antrophic and OpenAI does next.

ycui7 16 hours ago

OpenAntrophic and OpenOpenAI ?

kingo55 16 hours ago

Over Dario's dead body.

At least Sam have us a nice model last year with the gpt-oss family.

lebovic a day ago

I'm haven't found an announcement page, but there's a banner on the website announcing Qwen 3.8 and redirecting to this page.

Looks like they're previewing the model only on their subscription plan.

lebovic 18 hours ago

(This comment was originally on another merged post, and "this page" referred to https://www.qwencloud.com/pricing/token-plan)

trvz a day ago

It’s available in the iOS app (or was for me), both logged in and out.

Alifatisk a day ago

Is there an iOS app for using Qwen?!

trvz a day ago

maxrumpf 19 hours ago

> "compatible" instead of "comparable."

It boggles my mind how you can train a frontier model but not write a tweet without an obvious typo.

lardosaurusrex 19 hours ago

at the risk of upsetting a lot of people mentioning this but like are you really surprised?

if youre going to use ai for everything youre gonna start losing your edge as you focus less and less on what youre doing and this isnt me just talking out my ass, like... the front page here is peppered with study after study and blogpost after blogpost about how its overuse can come to the detriment of one's own abilities and skills.

coca cola had the ad with the magical truck that changed its design, shape and amount of tires it had and if nobody noticed that before releasing it then im not sure why anyone might think that the people peddling the LLMs would somehow be immune to this phenomenon

halJordan 18 hours ago

This is the sort of unmitigated pedantry thats the real problem. Everyone makes typos and errors of this nature. Feymann and Hemmingway both did. You deliberately chose this level of error multiple times when you chose to not put an apostrophe in youre and failed to capitalize proper nouns

Come off that high horse

lardosaurusrex 16 hours ago

whyenot 18 hours ago

The people training the model are almost certainly not the same ones writing tweets. I don't know why that is mind boggling. We all make typos, at least the humans among us do.

segmondy 19 hours ago

let's see your grammatical correct tweet in chinese.

rrhjm53270 19 hours ago

Well, I think Chinese doesn't really have any grammar most of the time.

pvorb 19 hours ago

Now you can be sure they write their announcements by hand. Doesn't really matter, does it?

aloknnikhil 19 hours ago

It's OK because I'm not prompting the person who tweeted for my usecases.

scotty79 12 hours ago

mistakes are at the moment a signature quality of human written texts, enjoy them

Bnjoroge 10 hours ago

i'll take that 100% over llm slop.

hodgehog11 a day ago

The "second only to Fable 5" comment is pretty telling here. I remember early on when a lot of naysayers were saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model, provided they continue to allow people to use it. It will be genuinely exciting when an open model is able to beat it.

InsideOutSanta a day ago

I wouldn't call it a moat, but I would call it a noticeably better model. Subjectively, for my own work, I would rate the top models Fable > K3 > Sol.

But it's not like Fable is so substantially better than the other two that I would be seriously impacted if I didn't have access to it anymore. All three are amazing models, and of the three, Fable is the only one that regularly triggers refusals.

hodgehog11 14 hours ago

It really does depend on your application. In my domain (math research), it is substantially better. Fable can solve really hard tasks with surprising consistency. It makes mistakes, and occasionally refuses, but honestly, at the top level, ideas are the currency and the rigor is the busywork. The other models cannot come close in this domain.

If you couple Fable's idea factory with Sol's rigor, you get a real game-changer. It puts the emphasis on top-level ideas, and nearly trivialises the intermediate layers.

porker a day ago

It's always fun to see what works for others, because for my work it'd have to be Sol > Fable. Fable makes too many mistakes.

Coordinating agents though? Fable any day.

neevans a day ago

tbh even if its better model due to lot of restrictions its not that useful than opus.

NitpickLawyer a day ago

100% this. There's currently this [1] submission that hasn't gained much attention, but is really important. In this [2] incident report from HuggingFace, they talk about detecting an attack and not being able to analyse the logs / IoC with API models because of guardrails. If not even highly regarded reputable companies can't sort out access to SotA models for blue team use, the raw capabilities don't matter. They're useless paperweights (hah!), and nothing else. Having to resort to open models is insane!

[1] - https://news.ycombinator.com/item?id=48965243

[2] - https://huggingface.co/blog/security-incident-july-2026

throwa356262 a day ago

hodgehog11 a day ago

Tell that to my colleagues. Despite Sol getting the attention, Fable is really starting to have an impact on mathematicians right now. It has unbelievable insights in a lot of cases that can rapidly speed up progress.

vitalyan8184 21 hours ago

their "genuine moat" is that mythos is the only super heavyweight model right now. it's always been possible to train a 10T model and get 10% more performance over a 1T model.

had mythos been just Opus 5, with the same size and price as the previous opuses, then yeah, that would be a tie-breaker. but it's not.

ferrouswheel a day ago

I dunno, I find Fable slops alot. Sol is my workhorse. Fable can be creative but isn't very good at doing work reliably (or without endlessly burning tokens).

XCSme a day ago

I am not even sure if Fable is as smart as they say, I can't get it to answer almost any question, it always refuses for "cyber-security" concerns...

reckless a day ago

5.6-sol would be a better comparison given it's general availability and usage allowances

hodgehog11 14 hours ago

They are likely assessing based on "raw intelligence" benchmarks, rather than agentic ones. Fable crushes in those, but that doesn't necessarily translate to microscopic rigor, which is what most people use these models for. You only see it when you ask really tough questions.

ferrouswheel a day ago

And sol is much more reliable as a agent for doing work. Fable sometimes just goes on wild flights of fancy.

epolanski 3 hours ago

> Like it or not, Anthropic have a genuine moat

I don't think yours and my definition of moat is the same, when I can switch from Fable to competing models and have similar results.

Havoc 13 hours ago

More like a puddle than a moat

scotty79 a day ago

> saying that Fable was barely an improvement on Opus. Like it or not, Anthropic have a genuine moat right now with that model,

What's more interesting is that Anthropic moat shrunk to just that model. There's zero reason to use any other model from Anthropic right now. And once they take Fable off subscription there will be zero reason to have Anthropic subscription.

nprateem 15 hours ago

Fable is still as dumb as a post. I ask it simple questions and it routinely gets things backwards, prioritises things that should be subordinate to others, etc.

An example: It just suggested that I shouldn't raise the price of my saas because it'd complicate the arithmetic if I did 0 -> $100k YT channel instead of sticking to $20 p/m.

It's just a complete moron, like all of them.

dgellow a day ago

That’s not a moat though

matheusmoreira 20 hours ago

What's the point of Fable if we can't use it? I get to prompt it like 5 times on my subscription before it gets cut off, and even then I'm constantly fighting the insufferable safety classifier.

I'll switch to OpenAI soon because of this. I also can't wait for the day it becomes feasible to run these awesome open weight models on my own hardware.

Alifatisk a day ago

> You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.

You can also try it out on Qwen chat, Its free.

LaurensBER a day ago

> With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.

That's a massive model!

The shift from "value" models to "intelligent, huge and slow" models coming from China is an interesting change in strategy.

My main issue with GLM 5.2 and Kimi 3 is that they're extremely token hungry and thus feel slow(er) to use.

hodgehog11 a day ago

Value models are always going to be there; you can always distill from a larger model. Having a really intelligent model, regardless of the size, is much better for building confidence in your brand. That is a big reason why the US companies are still hanging in there.

charcircuit a day ago

The shift isn't new. Kimi K2, a 1T model came out July last year. I am happy that more labs are following the trend as its important for competitive open models to exist.

selcuka a day ago

Also DeepSeek R1 was announced 1.5 years ago with ~0.7T parameters, which was a huge model back then.

chronogram a day ago

And DeepSeek has been making huge progress on efficiency, and publishing about, so they came with a 1.6T model that is both fast and cheap to run.

sinuhe69 a day ago

The title is misleading. The link led to a pricing page/token plan and not about the new QWen 3.8 model.

Schiendelman a day ago

The submission link is to the twitter announcement. The body just has a different link to pricing.

Alifatisk a day ago

I remember when they released Qwen 3.7 Plus and Max. These models behaved way different from all prior models, it became too verbose. It wrote multiple paragraphs just to answer my prompt instead of the usual concise and direct way responding to me. I didn't like that at all, and I know Gemini also had this behaviour with with the Flash series until I managed to reduce it a bit with personal instructions (in the settings on Gemini website).

I haven't tried Qwen 3.8 Max yet, looking forward to it. My hope is that its way less verbose. Another thing I experience with the Qwen models is that I do not trust their benchmark scores at all. Have anyone played with Qwen 3.8 Max and can share their experience? Which model it come close to? Sonnet 5? GLm-5? DS V4 Pro? Flash? Gemini 3.5 Flash?

2001zhaozhao 17 hours ago

Now there are not one, but two incredibly powerful open LLMs. I think this level of capability makes general prioritization / high level decision making doable with the right harness, and now everyone has hard-to-interrupt access to them (since these are open weights and someone in the world is going to run them). This world is going to get really weird soon, in both good and bad ways...

theplumber 17 hours ago

Let’s better download them fast before Dario is making a scene again!

kingo55 16 hours ago

If he bans them, they'll just show up as torrents. Or we'll download them from Chinese hugging face.

rcarmo a day ago

I do hope they provide optimized A3B quants--that's been the sweet spot for usable local inference for me, at least.

sieste a day ago

What is a "credit" and how does it translate to tokens for the different models?

xyzsparetimexyz a day ago

Its the currency of the future.

ahartmetz a day ago

Since Euro and Dollar values are reasonably close, you can call both of them credits, maybe

tclancy a day ago

whynotmaybe a day ago

They use it in Babylon5 (supposedly) in 2260!

jxmorris12 a day ago

Why did Qwen stop producing open models? They've gone from building the best open models ~1 year ago to producing like the 10th-best closed models. I don't understand this pivot at all.

Edit: I saw online they do in fact plan to release this openly at some point – x.com/Alibaba_Qwen/status/2078759124914098291

InsideOutSanta a day ago

They've announced that they're releasing the weights for a 2.4T model soon:

https://xcancel.com/Alibaba_Qwen/status/2078759124914098291

fragmede a day ago

It's not a pivot, giving away the weights was a marketing strategy that they don't need to keep up with.

bertili 17 hours ago

Wait.. the Qwen Max models have never been open-weight. But it sure sound like that's what they intend now?

"Qwen3.8 is launching and going open-weight soon! With a massive 2.4T parameters..."

kumanday 12 hours ago

I did some comparisons to Kimi K3... and found the best is combining the two! Model diversity is powerful.

https://trilogyai.substack.com/p/qwen-38-max-benchmark-how-i...

notnullorvoid a day ago

Always nice to see more open-weights in the heavy model class. I can only hope this trend continues, causing OpenAI and Anthropic to crash and burn.

blfr a day ago

As much as I dislike 'em, this sounds mean spirited. And Alibaba admits in this very tweet that Fable is next level (it is).

notnullorvoid 19 hours ago

I dislike their practices, but the main motivation for hoping they'll crash is that I think their immense overvaluation posses too much economic risk.

Hamuko 21 hours ago

I want as much misfortune as possible to befall OpenAI and Sam Altman after what they did to the memory market.

brap 21 hours ago

yeodev a day ago

Same, you can think of China whatever you want but they're really good at giving big tech a reality check when it comes to AI. We now got a pretty wide range of open-weight models (from DeepSeek and MiMo over to Kimi K3, Qwen 3.8 & GLM-5.2) and I think it's most important that there's a variance not only between quality / intelligence and also cost.

I mean even the cheapest option for Luna is still more expensive than anything DS or MiMo is offering right now and I think a new Ministral model would also hit hard there because we also need some variance in model sources, we can't rely only on the US and China.

glimshe 21 hours ago

"China" isn't giving anything. These are Chinese companies leveraging their best competitive strategy at the moment: competing on price.

derektank 21 hours ago

satvikpendem 17 hours ago

catigula 21 hours ago

Damaging American AI companies just means that China has the lead with closed models.

China isn’t an altruistic state. They’re an aggressor in many fields, economic and otherwise, and this one of them.

phs318u 3 hours ago

Off topic, but comment threads like this, where the early comments rabbit hole into something I didn’t come to read, really make me wish HN had collapsible comment controls. My phone screen nearly melted from the insane amount of swiping/scrolling to get to the part of the discussion about the actual LLM.

boutell 3 hours ago

I'm making good use of the next button on original replies, but it's true that walking back through the parent button is a long walk sometimes.

Elzair 19 hours ago

Has there been any news on open weighting Qwen Image 2.0 and WAN?

We are spoiled in the LLM segment, but I would love to see an open source competitor to Flux.2, etc.

antiloper a day ago

Does anyone have the privacy policy of their token plan available? Want to check if they retain/train on inputs/outputs.

moffkalast a day ago

Lol, lmao even.

Of course they train on literally everything they get their hands on, like everyone else. If you need privacy, that's what local models are for.

adamtaylor_13 a day ago

It's a flippant answer to a real question. Anthropic, OpenAI, and even Grok have "Don't train on my data" knobs.

Whether you trust them is different, but there ARE knobs on other hosted AI companies.

nextaccountic 10 hours ago

moffkalast a day ago

revolvingthrow a day ago

The few tests I ran were by no means comprehensive, but while kimi felt like the real deal qwen seems a bit of a benchmark princess.

dannyw a day ago

Qwen3.6 is still the best agentic open weight LLM around 30b params (Gemma isn’t very good at agentic execution).

I also find the model is a lot more predictable and less “glitchy” when made to think in Chinese. You can do this in the system prompt.

akazantsev a day ago

> Gemma isn’t very good at agentic execution

I had no issues with it for C++ development with https://pi.dev. I'm yet to try it with Zed Editor. I don't rely on agents too much. However, I used it on Chromium's codebase to research some functionalities, let's say for searching. Requests like: check my last commit and do the same for SetterA and SetterB; it also ran without any errors.

anana_ 19 hours ago

Apparently agentic performance in Gemma was improved recently: https://x.com/googlegemma/status/2077449152062247219

Too little too late imo

brunooliv a day ago

Qwen is so much better than GLM or Kimi that this makes me genuinely excited!!

Pesto a day ago

The bigger models are usually worse though, hopefully they nail it this time.

rhdunn a day ago

Does anyone know if they intend on releasing open source/weights variants for 3.8 or whether 3.6 was the last model they are/were doing that for?

rolls-reus a day ago

will be releasing weights per their tweet announcing the model https://xcancel.com/Alibaba_Qwen/status/2078759124914098291

netdur a day ago

The only problem I had with Qwen, fine tuning on Colab, it takes 31 t/s while Gemma 4 is around 9 t/s, otherwise, one of best local LLM

blmayer 13 hours ago

Looking forward for a 20ish billion parameter version

eurekin a day ago

With 3.6 27b, I just stopped changing local models and started tinkering with things on top (like mem0). Feels genuinely useful and more than a toy

androiddrew a day ago

I have only been using 3.6 27B for coding. Is mem0 for agents like Openclaw or Hermes? How are you using it?

eurekin a day ago

It's a mcp, so connects quite easily to agents. With mcpo, I also connected it to open-webui (which has better support for OpenAI style tools/functions). Used it in claude code with that mcp plugin set-up too. Only ever used it for managing homelab information, but it met initial expectations. 27b is a great model, if grounded. The query about physical hosts and routing... I haven't found a single hallucination (altough Codex 5.6 as a reviewer mentioned something was wrong with some parts, and those were exactly the never properly documented ones. Codex/gpt had extra knowledge, because it was the conversation I used to set it up).

sidcool 21 hours ago

These models are great, but what's the potential use? No small entity can run them.

alex43578 21 hours ago

China encourages/prioritizes their release because it directly competes with American companies closed models.

MichaelNolan a day ago

If 3.8 max goes open weight, what are the odds they retroactively open weight the earlier releases?

Gecko4072 a day ago

What would be the point? Kimi is kind of forcing their hand.

kennywinker a day ago

Well, if nothing else, posterity.

The_resa 9 hours ago

Open Source Era is comming

softwaredoug 17 hours ago

It feels like an inflection point of lost US leadership in technology? A year plus ago you would say while China led in green energy and manufacturing, at least the US was ahead in software - as demonstrated by the state of US AI models.

We could point at a lot of factors on the US side. From political paralysis / head-in-the-sand attitudes towards emerging tech like green energy. To something of disdain for workers that will be impacted by AI (creating a backlash). To education that continues to lag. Add to this so many other self-inflicted economic wounds from the current administration.

I don't know if its nearly as terminal, as say the UK after WW2. The US is still large, wealthy, and resource rich. Yet at a minimum the triumphalism about US leadership after Trump was elected by the tech elite feels silly in retrospect.

Something I also think about is how much stronger The West overall would be if instead of antagonizing allies, there was a single ecosystem working closer together.

pessimizer 17 hours ago

> We could point at a lot of factors on the US side.

It's just lack of antitrust enforcement.

China pours money into tons of different businesses in the same industry and lets them fight it out. The only US business model left now is to shut down (or collude with) competitors and raise prices while cutting costs. All they have to do is cut Congress (and regulators, and individual judges) in. We've financialized everything for the sake of scammers, rather that finance being used for the sake of getting cash to the most productive organizations. We've optimized for corruption.

If we hadn't let the stupidest people in the world buy up everything, and made doing nothing with it the most profitable option, China would have never have blown past us.

The US Supreme Court has explicitly legalized "tipping" politicians. That's the biggest sign of degeneracy that a government could possibly achieve.

https://en.wikipedia.org/wiki/Snyder_v._United_States

try-working a day ago

in case someone wants to speculate in why chinese labs open source their models: https://try.works/why-chinese-ai-labs-went-open-and-will-rem...

sbinnee a day ago

If it offers more than opencode go, the entry plan looks enticing

nullbio a day ago

I predict that no one will use this and everyone will use Kimi K3.

embedding-shape a day ago

I've been playing around with K3 a bunch, but the verbosity of the reasoning makes complete e2e agent work basically cost the same as other smaller models, and I'm not seeing a huge difference in quality, just a way longer e2e completion time.

sunaookami a day ago

Same problem with every chinese model currently, they overthink way too much and take too much tokens and time.

embedding-shape a day ago

EgregiousCube a day ago

rubslopes a day ago

Why? Price? If the reason is performance, I've been using non-frontier models for cheap, and they run great for my needs (GLM 5.2, DeepSeek v4 Pro).

rurban a day ago

We'll probably use it, but for images. Qwen is still the best for images

jadbox a day ago

What's the price difference?

corv a day ago

Who is behind this site? Is this another frontend to Alibaba or a reseller in Singapore?

khurs a day ago

Go China, screw America*

*within the scope of open models only

dannyw a day ago

I like my Apache 2.0 licensed Gemma, and NVIDIA’s Nemotrons are decent bases for finetuning or continued pretraining, esp thanks to good documentation and tooling.

Oh, and Mira’s thinking machines lab dropped Inkling, a ~1T open weight model too.

This isn’t US vs China. This is open vs closed.

jimbob45 21 hours ago

This isn’t open though. Promises to be open later aren’t worth anything, given what we’ve seen and heard from AI execs making promises in this industry.

ycui7 16 hours ago

khurs a day ago

It's Sunday morning so I'm allowed to be facetious!

baist0 a day ago

can i get "code instruct" version of this? i want 7B and 14B to launch on my hardware.

raised_hand 20 hours ago

interesting, when will this race end?

ernsheong a day ago

So are locally-runnable models frozen at Qwen 3.6 now :/

worldsavior a day ago

Everyone wanted open models that would challenge Opus and Codex, here, you got it.

ernsheong a day ago

We need better coding models that can run on local hardware, i.e. 128GB VRAM or less

zozbot234 a day ago

seanmcdirmid a day ago

tormeh a day ago

Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.

ch_sm a day ago

In my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet.

mark_l_watson a day ago

regularfry a day ago

androiddrew a day ago

I have been running 3.6 27b on a dual AMD r9700 setup using Opencode and Matt Pocock's skills workflow for writing Golang CLIs. It's decent, but won't win any awards on code architecture. I guess you can try to AGENTS.md the deficits but I am just exploring its raw Opencode experience right now. Much slower than an API but still 3x times faster than I can read. Tuning it in with a community chat template and a specific penalty for repeats was the sauce needed to get it to work. I can probably start loop daddying it now over the tickets Matt's flow creates.

So yeah, it's the best local model I've seen. I am going to try the Qwopus 3.6 fine tune soon with the same spec and tickets and compare the output of both.

SomeHacker44 a day ago

seanmcdirmid a day ago

I actually have long discussions with Gemini about this and have wound up download a bunch of different models for different things. There is no best, just fast but worse, slow but better, agentic or not, reasoning or not great at large contexts, better world knowledge, uncensored, etc…. It’s a bit daunting actually since there isn’t really a one size fits all model that you can just use for everything.

ernsheong a day ago

Yes it's between this and Gemma 4 31B which is much slower, but looks like it won't ever get an upgrade. I have to conclude that the MoE variants are unreliable, and MTP sometimes just can't get tricky formatting right.

dofm a day ago

SwellJoe a day ago

cmrdporcupine a day ago

hnfong a day ago

People have been able to run DeepSeek v4 flash with a high spec Mac.

schaefer a day ago

I flip flop between qwen 3.6 27b and qwen 3.6 35b 4b active.

But there’s also the quantization of DeepSeek v4 flash called dwarfstar

cmrdporcupine a day ago

Gemma4 models are arguably better. Or at least about the same.

atemerev a day ago

The best model you can run locally is Kimi K3, as long as you have the hardware. If "what model I can still run on a something resembling something I can put on desktop without separate electricity and cooling water inputs", then it is probably GLM 5.2 (can be run on e.g. Nvidia DGX Station workstation). As long as you have about $100k-$150k.

dofm a day ago

Maybe, maybe not. Qwen 3.6 27B is literally just three months old. Hard to predict. Maybe it just wasn't worth making a 3.7, and after all, the 27B release was after the Plus release.

sampton 19 hours ago

This is reminiscent of the operating system wars and browser wars. In the end there will only be 2 models that can survive. 1: give it away for free or 2: locked in with top notch hardware.

dartharva 20 hours ago

I very much appreciate the existence of these free models, but in my experience Qwen has too high of a tendency to confidently give the wrong answer as compared to other frontier models.

gxs a day ago

The open weights vs frontier models is reminding me more and more of the Linux vs Windows I grew up with (slashdot randomly popped into my head saying that)

I have a feeling this is the next…frontier of that fight

One can only hope it eventually does as well as Linux

Archit3ch a day ago

Obligatory "Does it answer security questions?".

mannanj 19 hours ago

It's an interesting time to be alive when your local models are supposedly the pinnacle of what a free nation is capable of, yet the ethicality of the companies is disliked and their models restrict and limit you so much you root for the models from a socialist/communist state. If it wasn't for the effectiveness of propaganda, tribalism and psyops in this scenario my words wouldn't even be controversial and would just be seen as a truthful observation.

dluan a day ago

waic go brr

Umair_khan2324 a day ago

nice

cadlernox a day ago

Nice

nwhnwh a day ago

Open what?

smnplk a day ago

Do this giant open-weight models have less active params and could be run on consumer hardware or no ?

adrian_b a day ago

Any open-weights model that has ever been published can be run on consumer hardware, even on a mini-PC.

The right question is which is the speed that can be achieved on a given hardware and whether it is high enough for the model to be useful.

Until now, the speeds reported for running big LLMs with the weights stored on SSDs have ranged from as low as a token every 10 seconds or so, to as high as a few tokens per second.

With open weights models that you host yourself, you are not constrained to use any single model, because that is the one for which you pay a subscription.

You can use many models, each for whatever it is more suitable. You can use frequently a small model with a high inference speed, but for some tasks you may actually save time with a better model, even if it is much slower.

In my opinion, even at 1 token per second a big model may be useful for some tasks.

apitman a day ago

I can only imagine what that does to the poor SSD

adrian_b 2 hours ago

Alpha3031 20 hours ago

vitorgrs a day ago

SVG's pelican https://gist.github.com/vitordelucca/521c2d63c9b852c622e7648...

Made on the website, so not sure if on the API there's more thinking options...

esrauch a day ago

I feel like the pelican test can't be relevant anymore; the whole point was to to something that wouldn't be in the training set at all and now it is?

rhdunn a day ago

A parachuting flamingo? An aardvark driving a bus? It should be easy to randomize the animal and the mode of transport (or vary it with animal playing a sport) to create images not in the training data.

onlyrealcuzzo 21 hours ago

How about an animated SVG of a pelican doing the Macarena, profile view, spinning to face the camera on the last beats?

LatencyKills a day ago

Agree. It was interesting/fun for a bit though.

joegibbs a day ago

What about an armadillo playing a piano? There are so many potential combinations It would say something if the pelican looked great but the armadillo looked terrible

rvz 20 hours ago

> I feel like the pelican test can't be relevant anymore;

It never was. The point of this "pelican test" was for performative reasons, or just for attention of the joke.

It is like trying to test whether if an adult elephant could actually climb up a tree and reporting that some elephants are slightly better at doing that than others while also reporting at the same time that they are all bad at tree climbing anyway.

This is an example of testing for the sake of testing. The "pelican test" tests for nothing.

cakbeslik a day ago

Using QWEN models since 2.5. I never used the chat properly but as an API I can say they're quite good, especially when you compare with OpenAI models. Cheaper and almost same level. I will try this now also.

comandillos a day ago

Just imagine Anthropic making Opus open-weights now for the sake of trolling everyone. Wouldn't surprise me at this point xD

hodgehog11 a day ago

That would go against everything that Dario believes in (note that I refer to the CEO and not the company; the staff at Anthropic are not so ridiculous). He believes in Anthropic being the sole arbiter of the forefront of this technology, because it is all too dangerous in the hands of anyone else.

anon373839 a day ago

I’ve seen no evidence that he believes in anything. He comes off as just another slimy would-be monopolist to me.

cyanydeez a day ago

embedding-shape a day ago

That OpenAI releases Sol as downloadable weights feels way more likely than Anthropic releasing even the tiniest of models for download.

ferrouswheel a day ago

Has Anthropic released anything at all, ever?

kingstnap 19 hours ago

Anthropic has a lot of interpretability work, but they are extremely defensive about everything else. Dario doesn't really believe in releasing anything. Despite how he acts in terms of some sort of highly principle driven saviour (machines of loving grace) he clearly is more business minded then anything else.

For example they don't even tell you anything about the tokenization. They even do random chunking and padding to avoid leaking the token strings in the streaming api after it got reverse engineered. (See: https://spylab.ai/blog/claude-tokenizer/)

Kuxe 17 hours ago

Not voluntarily

ferrouswheel 6 hours ago

Alifatisk a day ago

Model Context Protocol

ferrouswheel 6 hours ago

anonym29 a day ago

safety datasets and a lot of safety related research