Muse Spark 1.3 (developer.meta.com)

437 points by bvaldivielso 8 hours ago

simonw 8 hours ago

  llm -m meta-ai/muse-spark-1.3 "Generate an SVG of a pelican riding a bicycle"
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

4.2266 cents, 38 seconds.

For comparison here's Muse Spark 1.2, which animated it without me asking it to: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The 1.3 one is definitely better - better bicycle frame, better wing, better pelican hat.

UPDATE: Here's another one with five pelicans for each of the five Muse Spark 1.3 reasoning levels: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

The most expensive was reasoning level xhigh - 7.5 cents, 1m34s.

And I ran five pelicans at all reasoning levels for 1.2 as well, here: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

wewewedxfgdf 5 hours ago

I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle.

Aced it, got the job as a senior software engineer.

The interviewers afterwards said "it is SO refreshing to find a software developer who actually knows how to code - never seen such a high performance focused, well built pelican on a bike - you have the skills we need".

labrador 3 hours ago

I interviewed as a software developer at LinkedIn. The interviewer asked me to demonstrate my prompting skills, so I had AI write an article about what the recent death of my father taught me about B2B SaaS. Reading it brought tears to his eyes so he hired me on the spot.

latentsea 2 hours ago

rattray 2 hours ago

nycdatasci 2 hours ago

hunterpayne 2 hours ago

fuddle 5 hours ago

I also aced my interview by focussing on pelicancode problems, instead of leetcode problems.

salutis an hour ago

"I was interviewed for a job as a software developer last week and they asked me to draw a picture of a pelican riding a bicycle. Aced it, got the job as a senior software engineer."

That is the best joke I have heard this year. Ready for a stand-up comedy special. Or a song. Superb!

sroussey 3 hours ago

You should post your source code you wrote here… ;)

UltraSane 3 hours ago

If you could actually hand write SVG code on the spot that looked like a realistic pelican riding a bike I would want to hire you for SOMETHING.

treebeard901 15 minutes ago

drusepth 8 hours ago

Is there a reason these pelicans always have roughly the same composition (side-view, 2d, biking right, flat ground beneath, etc)? I don't see any of that detailed in the prompt, yet they all seem to generate roughly the same image of differing quality.

bodeadly 6 minutes ago

Yes. It's because you are asking it to generate an image of a pelican riding a bicycle. If someone asked you to draw a pelican riding a bycycle, would you interpret that to mean using 3d photorealism? LLMs follow conventions. The convention for an animal riding a bike is to create a childish 2d line drawing.

vunderba 8 hours ago

The more generic your prompt, the more generic the response. It's a regression to the "mean" of the training data aka GIGO for AI.

It's like when you ask your average person off the street to draw a house - it'll almost always be square with a triangle roof, one door, and two windows.

In the pelican/bike example, it's probably a bit of a self-perpetuating snowball too. If the earliest examples were bike left-to-right, flat ground, etc. then they are also being scraped up in future LLMs.

werdnapk 4 hours ago

polyterative 8 hours ago

BeetleB 6 hours ago

Not when rendered via POV-Ray:

https://blog.nawaz.org/posts/2025/Oct/pelican-on-a-bike-rayt...

I plan to update it with more pelicans from all the models released since.

(Spoiler alert: They haven't improved much since then).

murkt an hour ago

xhrpost 6 hours ago

fc417fc802 an hour ago

simonw 8 hours ago

It's really interesting, isn't it? They almost always cycle from left to right - but I have had a few which cycle in the other direction.

The 2D / flat ground feels reasonable for a SVG, which implies a vector illustration.

m12k 7 hours ago

johntb86 7 hours ago

piker 8 hours ago

porphyra 6 hours ago

The canonical view of a bicycle is facing right. Usually, people want to draw/photograph/depict the side of the bicycle with the running gear, which is on the right side of the frame for historical reasons.

postalcoder 8 hours ago

Search Google Images for "bicycle". Almost all bicycle product shots are staged the same way: side view, going left-to-right. It makes sense to me that given that skew in the training data, the model grounds itself in the bicycle.

daemonologist 8 hours ago

threetonesun 7 hours ago

__MatrixMan__ 5 hours ago

The thing that distinguishes pelicans from other birds does so most strongly in profile. If you're looking straight at one, the throat pouch would be hidden by the beak.

I bet if it instead had something to do with black widow spiders we'd find that we're most often looking at the bottom of the spider's abdomen, regardless of whatever non-spider-like activity is supplied.

elfly 2 hours ago

well it is svg, it is doing it from circles and lines as primitives, it wants to do it simply and kind of builds the whole thing hierarchically. Making it 3d is way more complicated (as the POV example shows) and the prompt doesn't say 3d anyway

SV_BubbleTime 3 hours ago

I’m a firm believer in pelicanmaxxing.

They’re all so close in proportions.

optimalsolver 8 hours ago

Sun is missing a few rays and not wearing sunglasses.

ModernMech 8 hours ago

Yes, I do a thing where I ask the machine to generate responses in the form of a lizard talking to a cat. The lizard is always a green gecko and the cat is always orange, which I never specify.

reaperducer 7 hours ago

Is there a reason these pelicans always have roughly the same composition

Because they're computers. They don't have an imagination and the ability to create things from whole cloth the way humans do.

Much like a mother pelican, they regurgitate what they've been fed.

drob518 7 hours ago

Simon, at this point I really wonder if teams aren’t gaming this. You should pick a random animal doing a random thing every time.

gpt5 3 hours ago

We should just consider the pelican bench as saturated and mostly meaningless.

tintor 7 hours ago

Did any LLM so far draw pelican knees correctly and have them bend in opposite direction from human knees? Knees of many animals bend opposite to humans.

Did any LLM draw the front bicycle wheel correctly? ie. center of front wheel slightly AHEAD of steering wheel axis. This is done for bicycle stability.

TiredOfLife 7 hours ago

jonahx 8 hours ago

If you have a grading rubric, huge points off for adding arms instead of using the wings as arms!

Fergusonb 8 hours ago

I think it's hilarious that this detail is enough for me to dismiss looking into the model, but here we are, and it is.

hollowturtle 7 hours ago

Is there any point anymore regarding this svg test? I would not be surprised if in the training they're fine tuned for this task too

tintor 7 hours ago

It would be very embarrassing for any lab to benchmaxx the pelican on bicycle svg prompt, since it would be very easy to detect it by varying the prompt.

nojs 5 hours ago

BeetleB 6 hours ago

cheesecakegood 4 hours ago

It also works as extremely effective engagement farming, for lack of a better phrase

pavs 4 hours ago

FYI, your renderer breaks with error "git api access error 403", rate limiting error from git, when using cloudflare vpn.

I am guessing its not super common, but it happens just so you know.

m00dy 2 hours ago

I see no point having these pelicans used for anything related model qualification.

jonplackett 8 hours ago

Has any ab tried to game this yet and just made the most amazing pelican by hand and always reply with that?

ipsum2 6 hours ago

All of the links show "Error: Gist API returned 403".

leumon 5 hours ago

next, try: "generate an svg of a human hand". this is a prompt where many models fail imo.

jttnr 7 hours ago

I wonder, given Simons reputation in AI benchmarking, whether model providers try to train or tweak their models to perform better at drawing bicycles and pelicans?

andytratt 4 hours ago

excellent thread

tomrod 8 hours ago

What does the mean pelican look like at this point?

Also 3X token use vs. 1.2

_puk 7 hours ago

Red eyes and a tattoo?

EugeneOZ 7 hours ago

Absolutely BRUTAL! :)

Thank you for doing this, I love your benchmark the most!

0xbadcafebee 6 hours ago

For all the comments of "I'm sure they're fine-tuning for pelicans": https://dylancastillo.co/posts/pelicanmaxxing.html

  "Sorry, HN haters, but there’s little evidence that AI labs are pelicanmaxxing.
   Or at least they’re not doing it in a plainly obvious manner."

jmkni 8 hours ago

lol

Definitely an upgrade over 1.2

superfrank 8 hours ago

I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap and was actually really pleasantly surprised with it. It's not a frontier model by any means, but for work that didn't require a top of the line model, I really enjoyed using it.

I'm anthropomorphizing it a bit, but it felt like it knew its weaknesses and didn't try to impose it's opinions on me. What I mean by that is that it did what I told it and if there was something unexpected in the code that it put out it was often because I gave it ambiguous or conflicting instructions. It didn't try to go above and beyond and just acted like a tool, which is what I want from a coding agent 90%+ of the time. I also felt that it did a much better job of following established patterns in my code than many of the other current models do. I'm a huge fan of OpenAI's models and Spark 1.2 is what I expected 5.6 Luna to be.

I'm curious and a little excited to use 1.3, but honestly a little worried that as Meta pushes for better benchmarks that Spark will start to fall into the trap of trying to be "helpful" in ways I don't want it to be.

Tangential, but when I first started using Spark 1.2, it made me realize how much I miss 5.3 Codex. That model was the peak of coding models, IMO, in that it knew how to write good code, but didn't try to overstep or be "helpful" in unexpected ways. That got me thinking about how the major labs seem to be stepping away from coding focused models toward more general purpose ones and how I can't help but feel like that's a mistake.

MangoCoffee 7 hours ago

>I started using Spark 1.2 for development because if you're willing to let Meta train on your data it was dirt cheap

its free on opencode and i use it for personal projects. most of my personal projects are AI generated since its personal projects. nothing important are on them. it is hilarious if Meta is training their AI model with AI generated code.

CGamesPlay 4 hours ago

The useful training data is when you clarify your intent, when you tell the model a different approach would be better, when you consistently refactor towards Y and away from X, and so on. The training data isn’t the code, it’s the session transcript. (Anthropic would call this a “distillation attack” against their model, but in this case the model is you!)

superfrank 5 hours ago

Funny. I use it through Opencode Go which gives more use than I can use, but didn't realize it was actually free on Zen. Will switch to that I guess

KptMarchewa 7 hours ago

I would imagine your interactions with it are more important than the output.

dakolli 6 hours ago

Every lab trains their models with AI generated code at this point.

dcl 4 hours ago

TiredOfLife 6 hours ago

Training on ai generated content is how the models got a big jump in capability

sejje 7 hours ago

If it's a mistake, it should course-correct.

I agree that some of the smarter models are actually worse. I hope they take a model that's good enough--there are many--and just try to get it chatjimmy.ai speed.

I have to think that's the future, somehow, and I'm really excited about it.

bertili 8 hours ago

DeepSWE scores 75.4 - that's the best score so far. And it's crazy cheap! Google held the top a few hours today with Gemini 3.8 Flash, but now second to Spark 1.3. All this competition will drive prices down!

notatoad 2 hours ago

when are we going to stop pretending these benchmarks have any meaning?

anybody who's used these models knows that their real-world software engineering performance has no relation to the ranking on deepSWE.

caconym_ 2 hours ago

+1. I've used the recent Gemini Flash models and I've used Opus 5, and the latter makes the former look like a box of broken crayons. Unless Flash 3.8 and/or this Muse Spark model are a much bigger deal than people seem to think, I will eat my hat if either one can come close to Opus 5 in actual real life "long-horizon software engineering" tasks.

(I'm not happy about the above being true, but it's the reality I seem to inhabit.)

jdm2212 30 minutes ago

WASDx 8 hours ago

With the contributor pricing being more than 10x cheaper than the standard, that would make it best and cheapest on the DeepSWE leaderboard! It feels fast in my experience too. LLMs keep improving at an insane pace.

dakolli 6 hours ago

and they're ultimately tools strictly to replace you and your labor, they can't/won't cure cancer or make your life better. Your life will get worse and worse in every aspect until they extract maximum value from all of our lives with this technology through every avenue possible. Not sure why you guys are so excited about these developments.

This technology is strictly an extractive parasite on the world. Use it, but don't be excited.

yipinwong 4 minutes ago

comicjk 5 hours ago

nl 5 hours ago

cycrutchfield an hour ago

switchbak 2 hours ago

skybrian 6 hours ago

israrkhan 2 hours ago

Gemini 3.8 flash has better rates. $0.75 per million input tokens and $3.75 per million output tokens.

Compare that to Muse spark 1.3

$1.25/M input, $4.25/M output (without data sharing) $0.10/M input, $0.20/M output (with data sharing)

It is dirt cheap, but only if you are willing to share your data with meta and allow them to use it for improving their models and products.

cbg0 8 hours ago

But is the score really reflective of the quality or are both models benchmaxxing?

gpt5 3 hours ago

Both versions of DeepSWE (1.0 and 1.1) are likely not that meaningful anymore. Whether through models progression or through contamination.

bermudi 7 hours ago

Muse 1.2 wrote a terrible "smart summaries" extension for my pi setup. It was sending every single steamed chunk for summarization instead of waiting for the full CMD.

This is an error I would expect from sonnet 4, not a model that was supposedly just a few points behind sol.

dominotw 8 hours ago

how much of it is from reallocation of staff to ai training and labeling

IIIIIllIIII 8 minutes ago

Im a caveman writing c/cpp. Last time ms1.2 was even worth than DeepSeek v4f preview on internal benchmark. It just feels like extremely over fitting on certain paths.

Lucasoato 8 hours ago

A model that (at least in benchmarks) is getting closer to SOTA. A clear separation between what’s used to improve their products and what’s not (at least this is what they claim).

Good job Meta! Seriously. This is almost making me forget about the 18B$ lawsuit for children social media addiction.

dbbk 7 hours ago

How is it not SOTA? It's beating 5.6 Sol.

wrsh07 4 hours ago

It's somewhat useful to note just for your own timelines that Fable was reportedly trained in February. I'm not sure when mythos 5.1 finished training, but muse spark 1.3 almost certainly finished more recently than that.

This doesn't mean it's not one of the best models available (clearly it is), but that table didn't compare Fable/mythos (unless I missed it?) and OpenAI will be releasing a much more recently trained model (Astra) any day.

So you shouldn't think "wow, Facebook has caught up"

You should think, "wow, Facebook is less than 6 months behind the frontier" and that they're actually creating good models which is going to be good in many ways (price for customers, for one!)

There are downsides too, but I'll discuss those separately somewhere

ctolsen 6 hours ago

You gotta keep up. Fable 5.1 came out yesterday and is better so anything else is to be treated as garbage now.

pqdbr 6 hours ago

neuronic 5 hours ago

jmward01 6 hours ago

muse-spark-1.3-contributor. Say what you want and Meta, changing the pricing to explicitly say 'we train on this and value it this much' is what every model provider should do. As a side note, it is now completely obvious how much stealing my tokens for training is worth to model providers. I avoid/pay extra/try my best to make sure I am not getting trained on but it seems like it keeps popping up that I missed a setting somewhere. This is the first quantifiable number I have seen out there from a model provider. Maybe it can help in lawsuits to quantify the damages for copyright/other things?

popularonion 18 minutes ago

This has been my hunch for a while about all the discourse of "OpenAI/Anthropic subscription pricing is unsustainable!!"

We understand theoretically they're taking our data, but yeah, that data is vital to the entire business plan of all these companies and WAY more valuable than people are giving credit for.

I checked up on Mistral recently and saw their Claude-alike coding harness is using GLM now, whatever it takes to keep users on their platform and feeding them data.

apodolny 6 hours ago

I like the approach of providing a discounted version of the API that is used to train vs. the full price version. Seems reasonable and transparent.

Gecko4072 8 hours ago

Used Muse Spark 1.2 and was not impressed at all. Fast and cheap but even GPT 5.6 Terra felt much more capable. Also not really looking to support a company that was just forced to pay $18B for mental health damages.

WASDx 8 hours ago

I'm party using 1.2 to reverse engineer and re-implement an old game binary and it has been quite good and fast. The contributor pricing is very attractive, excited to try 1.3 and see if I feel a difference. 1.2 can get stuck outputting similar sounding thought summaries with no apparent progress when asked to solve bugs. Then I've switched to GLM-5.3-Flash which for this use case has been clearly better at finding suspected causes and following tracks.

geoffbp 3 minutes ago

> /taste: an anti-slop filter: a flat checklist of visual defaults not to use, so generated UI stops looking machine-made.

This is interesting

ydna404 30 minutes ago

For folks who are impressed with costs, why does it matter to you? Is subscriptions not a thing? I may be missing something but only companies should really care about this I would think?

MitziMoto 26 minutes ago

Some of us own and run companies? Cost per performance is a huge deal.

7734128 8 hours ago

Practically free for "contributors" at 0.2 usd/mtok. That's going to be hard to say no to for hobbyists.

wrsh07 4 hours ago

Price segmentation at its finest

0xbadcafebee 6 hours ago

I'm wondering whether anyone has yet extracted AWS keys from a model trained on user input. Because users are definitely feeding secrets into these "contributor" models

userbinator 10 minutes ago

If my experience with image generation is any indication, unless AWS keys are somehow extremely prevalent in the training data, you may get something that looks like one, but it definitely won't be valid.

HDBaseT 5 hours ago

A small number of inputs in a large dataset can poison training data pretty drastically. Anthropic wrote a good article about it a while back [0]. This should mean its possible to pull back that information fairly easily.

It is hard to not feed it "secrets" too. Models will see path names, read compose files, etc. Of course you can configure things to not leak this type of information, but its not default in most harnesses and isn't 100% sufficient anyways.

[0] https://www.anthropic.com/research/small-samples-poison

owaiswiz 5 hours ago

doesn't mean the raw text goes into training. they most likely have a pipeline to clean out any secrets before they train on it?

HDBaseT an hour ago

wxw 8 hours ago

“contributor” pricing at $0.10/$0.20 is crazy cheap if it’s measuring up to Sol.

Definitely shows how important a user data flywheel is for RL and model improvement.

majerep 8 hours ago

The previous version was, in my experience, the best free model available on OpenCode. It's been very good at simple/moderate tasks where I am precise in my ask and it doesn't need to make a ton of undefined assumptions. Hopefully this new version is also available on opencode for free.

jumploops 7 hours ago

The "contributor" pricing is the standout here at a ~20x discount, if you allow training on your data.

The model seems on par with Sol and Opus 5 on paper (admittedly on some older/saturated benchmarks, but very competitive for $).

Stats:

1M context, $0.10 input/$0.002 cached, $0.20 output (Mtok)

HDBaseT 5 hours ago

Not to mention, this is hyper competitive against even Chinese providers given its multi-modal support.

Muse Spark 1.3 supports Text, Image, Video, File, Audio inputs. We've only started to see models from China include image and video inputs recently.

a012 3 hours ago

Muse Spark may be competitive in capabilities but it’s not for serious works since Meta trains on your prompts so no ZDR, in contrast Chinese provider like Z.AI promises ZDR which is more attractive to big corps.

2001zhaozhao 7 hours ago

I have a feeling that Meta is not gonna like what people actually use the contributor model for lol.

(It's probably going to be a bunch of repetitive batch jobs like web search that have no training value)

winstonp 7 hours ago

It's the perfect model for open-source work because it's gonna end up in the training data anyway

hadlock 6 hours ago

There's a lot of value in agentic loop tool failure + recovery training data

dbbk 7 hours ago

Web Search doesn't have a discount on contributor pricing

cnxhk 7 hours ago

keyle 3 hours ago

I am very impressed by this model so far. It's faaast and it seems to be just intelligent enough to do really well. It's UI work (simple python UI) is very clean and functional. The UX was 'there'.

coolcoder613 4 hours ago

I have not tried Muse Spark for code, but I've been using it for a while to write Latin. I find it's one of the best at it, alongside Gemini. For example, I've recently been using it to translate the subtitles of the show I'm watching into Latin, to provide me with a bit more input. (I'm learning Latin, for context)

dcl 4 hours ago

Very keen to try this after using Claude Code over the last few months. Should I just point Claude Code to Muse Spark endpoint (because I'm familiar with Code)? What do people think of Muse Code or other coding agent harnesses?

alexboehm 4 hours ago

Just try opencode, it comes with 1.3 contributor free.

dcl 4 hours ago

Well thats very interesting. Thank you. Will be interesting to see how hard/easy it is to translate my Claude skills, loop design, etc to the new harness.

This kind of raises another question to me regarding the coding benchmarks, how much of it is model versus harness?

ryanschaefer 5 hours ago

For all of the comments about training: I thought that subscription plans for other models allow the same. Am I mistaken?

Aurornis 5 hours ago

It's a toggle. Some will automatically enable it and you have to turn it off. People who rapidly click through setup flows can miss it and leave it enabled.

israrkhan 2 hours ago

it seems like gemini 3.8 flash is more capable and cheaper. The only reason i would use this is if i was willing to share my data with meta, and allow them to train on my data. In that case it becomes dirt cheap.

fibonacci112358 7 hours ago

Is everyone rushing to launch something before Astra tomorrow?

maciejgryka 6 hours ago

Does anyone know what the license for this model is? Specifically any word on restrictions about what it can be used for?

gehsty 6 hours ago

As a product, would developers switch to a meta model/harness? I don’t think so.

Only way I see is if it becomes the new SOTA / frontier, does anyone think Meta will surpass Anthropic or OpenAI?

I still can’t get my head around why language models are an existential threat to Meta - they own the platforms people watch adds on?

phyrex 6 hours ago

Meta also has 50k engineers. Not to mention that tons of meta infrastructure - including ads! - use AI. Would you want that sort of business be this dependent on someone else?

souvlakee 8 hours ago

Why they didn't use LLM to create html table instead of https://lookaside.fbsbx.com/elementpath/media/?media_id=1048...?

yanjunnf 3 hours ago

It's true that there hasn't been any meta news about LLM for a while now

esafak 2 hours ago

Funny how quickly Meta caught up after Lecun left.

wkcheng 4 hours ago

How do people actually use this? Do they use it through some sort of subscription plan, or via OpenRouter?

dv35z 4 hours ago

You can check out Muse Spark 1.3 by using OpenCode (https://opencode.ai/ - open-source AI / coding harness). There's a terminal version and a GUI / desktop version. Good luck!

scotty79 8 hours ago

Is the fact that everybody almost catches up with the frontier a sign that we are entering a new region of sigmoid curve?

schopra909 8 hours ago

Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up.

Moreover,I think it’s impossible to know if you’re hitting a portion of the sigmoid, because there will often be an idea that changes the trajectory altogether.

In 2024, there was a ton of talk about the plateau. Reasoning was an iteration on chain of thought, but it didn’t really work. Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data. That small iteration catches the eye of OpenAI and Anthropic, turns out to be way more important than even DeepSeek could have ever expected when it comes to improving LLMs for coding, and last 18 months have been an exercise on riding that insight to the nth degree.

That one small iteration brought us a lot of progress. Now we’re seemingly exhausting the impact of that one insight, but there may be another soon enough.

stymaar 8 hours ago

> Deepseek proposes RLVR as a way to get around the lack of $ they have to produce human reasoning trace data.

What was the difference between what deepseek did for R1 and what OpenAI did for o1?

npn 29 minutes ago

Philpax 3 hours ago

o1 was first, and Anthropic were doing a bit of it; DeepSeek brought it to the masses, but did not invent it.

schopra909 38 minutes ago

refulgentis 7 hours ago

I don’t know why people think DeepSeek did reasoning models / RLVR before OpenAI, there was a gap of months.

danielmarkbruce 6 hours ago

Even if all the big ideas are gone and we are entering a new part of the curve, there is still an enormous amount of improvement possible. Just iterating on data mix/quality etc, training pipelines, reward functions, specific ways of reasoning (which i guess is mostly just data still) for the next 20 years will yield a looooot. And that's just the models. The harnesses/application layers/whateveritgetscallednext space has 20 years of progress to make.

samuelknight 8 hours ago

Meta has an enormous amount of compute. They are either going use it making and inferencing models or they are going to sell their excess capacity to model providers. Zuck had to completely rebuild his AI team after the Llama 4 launch mess.

ipsum2 6 hours ago

Yes. It's really up to OpenAI/Anthropic to release a new paradigm to shift the curve now, before everyone catches up entirely.

gdiamos 4 hours ago

I think it means that we should be aiming further ahead

redox99 8 hours ago

No because the frontier keeps advancing very fast.

dominotw 8 hours ago

meta fails at everything yet is frontier on this one

geooff_ 8 hours ago

Could this be best intelligence / $ if you're willing to let zuck digest your data?

HDBaseT 4 hours ago

By default, even without the training endpoint the pricing is pretty competitive, especially against Opus and Fable. [1] The 'muse-spark-1.3-contributor' endpoint is by far the cheapest, significantly cheaper per M than ChatGPT Luna, significantly smarter than Luna too.

This price/intelligence beats even legacy DeepSeek V4 Flash pricing.

[1] https://artificialanalysis.ai/#total-cost-tabs

oofbey 5 hours ago

Yeah. Super icky. But this might be the first time in Zuck’s life he’s being honest about the business model.

mromanuk 8 hours ago

I didn't like 1.2, It make some mistakes in a web app, so I quickly went back to Claude, Kimi K3 or Deepseek V4. Hope this one can clear agentic development, because Muse Spark models are fast and cheap.

LZ_Khan 6 hours ago

Ha, even with monitoring engineers keystrokes and mouse movements not SotA on OSWorld.

finnjohnsen2 8 hours ago

So one model is "Not used to improve our products" and is 10-20 times more expensive to the "Used to improve our products"-model.

Given this is Meta, my immediate assumptions that one is cheap because it lets me "be the product". I know I'm rushing to conclusions but there is zero trust here. The brain will do its thing. And the wording here is giving the brains a lot of wiggle room.

duplessitous 7 hours ago

What is the confusion? They directly state that you are the product if you use their discounted offering. It isn't an assumption that should lead you to this, it is Meta's very direct communication that should lead you to this

Jcampuzano2 8 hours ago

I'm confused what your surprise is here. It's plain and simple right to the point wording.

I don't see the wiggle room at all.

thefreeman 8 hours ago

aren't they explicitly saying this with both their pricing and their wording? I'm not sure what you are alluding to?

whimsicalism 8 hours ago

the meaning is pretty obvious - they want to train on your chats & tasks and are willing to subsidize for the privilege of doing so.

r_lee 3 hours ago

they're doing the same thing as DeepSeek

zhoBEENG 7 hours ago

Would it help you understand if they were labelled "For Dumb Fucks" and "For Everyone Else"?

warkdarrior 6 hours ago

Privacy is not free. They make it quite clear that they charge more if you don't want your data used by Meta.

IshKebab 7 hours ago

I think it's more that the "not used to improve our models" is expensive because companies need that. It's simple price differentiation.

In other words, it's not that Meta really wants your data and they're willing to pay top dollar for it. It's that companies really don't want Meta to have their data and they're willing to pay top dollar for that.

bigyabai 8 hours ago

Given OpenAI and Anthropic's behavior, do you really expect them to be singled out for this practice? Zero trust has been in "LGTM" territory for years now. Meta's bet against people taking a principled stance arguably paid off great.

ChrisArchitect 8 hours ago

sunaookami 7 hours ago

>Previously available reasoning modes are available today with max reasoning coming shortly after we finish additional safety testing

Lmao. And their benchmark table only shows max reasoning.

m00dy 2 hours ago

$META has everything it needs, great team, great models coming out, great infrastructure (GPUs), great userbase and distribution channels. $META is underrated.

mmastrac 6 hours ago

Any idea what size this is?

anjel 5 hours ago

Not mentioned in pricing: Surveillance costs of using Muse Spark

lostmsu 7 hours ago

What a day. OpenAI is behind basically all major competitors - at least for a some amount of time.

Bolwin 4 hours ago

I highly doubt it's behind in practice, except for Anthropic

tinyhouse 8 hours ago

I had no idea Meta has a coding agent harness. Does anyone have experience with it and can comment? The 1.3 contributor prices look very attractive. I'll probably start using their API if performance is good and the API is reliable with decent rate limits.

meric_ 7 hours ago

You should use their harness. They trained it on multiple harnesses but have specifically optimized it for their harness. Cline also did an independent experiment w spark 1.2 where using the native harness makes it use fewer tokens / turns to accomplish tasks

tinyhouse 15 minutes ago

Thanks. Just downloaded and pretty impressed so far. It's fast and nice to work with.

dcl 4 hours ago

Any more info on this?

meric_ 2 hours ago

frozenseven 8 hours ago

This should probably be primary:

https://news.ycombinator.com/item?id=49541149

IshKebab 7 hours ago

Lol "not used to improve our models" is AI's enterprise SSO.

r_lee 2 hours ago

that'd be ZDR, the one you need to beg from their Sales teams with $$$

tyre 8 hours ago

Meta is one of those companies where, if there is anything remotely comparable, I'm happy to pay more to not use them. They've had a profoundly negative impact on society and Zuckerberg is not who I want controlling the future at the top of AI.

I feel the same about Grok w/ Elon. I will pay extra to use someone else.

I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.

And, yeah, I wouldn't trust sama to watch my bag while I went to the bathroom.

biddit 6 hours ago

Strong disagree with the Anthropic being good at all part. This is not defending anyone else, but…

Anthropic leadership repeatedly presents themselves as uniquely morally qualified to steward agi and decide how humanity should get access to it. Yet they have repeatedly failed basic morality tests.

Pirating books for financial gain. The newer Sony/Warner music case shows this is pattern behavior.

Aggressively scraping other people's works, despite the authors' requests not to do so.

Then applying massive usage restrictions on their own work.

And probably the most disqualifying is backing away from their own hard AI safety commitments.

nostromo 6 hours ago

It makes me sad that people don’t see right through Anthropic’s gambit.

They want to position AI as an insurmountable threat in order to regulate away any future competitors. They’re trying to speedrun regulatory capture.

sscaryterry 6 hours ago

usef- 6 hours ago

Which safety commitments did they back away from? My understanding is that they believe safety can only be researched from the frontier, and so they're trying to be pragmatic to stay near the frontier (and viable) in their choices.

From what I know, the "books3" dataset was normalised in the LLM and research ecosystem, where collected datasets were seen as valid to train on and/or fair use. I'm not sure any of the major frontier companies are free from that, if we don't believe it was fair use.

I do think most of their choices are explainable by "they just believe in agi risk". You truly wouldn't want non-agi-pilled companies to train on your data and approach the frontier if you were worried. You might slightly hurt your own business with safety filters (that no one else does) if you were worried. They are less worried about other "moral" decisions like "sharing" if they conflict with AGI: the research they still share is all of their safety research.

This definitely doesn't make them "good", but they do seem fairly "consistent". Most of these issues were talked about publicly by the founders long before Anthropic was founded and/or the AI race+money appeared.

8note 5 hours ago

dofm 6 hours ago

> Anthropic leadership

Which one? The main bit that reports to Daniela Amodei, or the little comfort blanket cabinet around Dario and his "chief of staff"?

There is a leadership branch that can pretend to be morally qualified and aware and to think about the big picture and ethics.

It is at least somewhat remote from the bit that is doing the actual business things.

janalsncm 5 hours ago

I will give Anthropic credit for standing up against the department of war. The bar is incredibly low, but not doing domestic surveillance and not creating autonomous weapons are laudable.

That doesn’t mean I like them pirating books and being shady about tokens and paternalistic “safety”

marcuschong 5 hours ago

It's hard for me to see much difference between Amodei and Sama. My guess is they're both savvy SV CEOs who will bend their message, alliances and principles pretty far if that's what it takes to get ahead. Musk and Zuck feel like something else entirely, with all the reactionary imagery, populist bullshit and the societal damage around their platforms.

vovavili 6 hours ago

It's almost like running a trillion-dollar business with neck-to-neck competition against other frontier labs and even state-sponsored efforts requires some ethical trade-off.

ACCount37 6 hours ago

Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense.

It makes a lot more sense than having to work around copyright by scanning out physical books. Unfortunately, one was ruled legal and the other was not.

jwitthuhn 7 hours ago

Dario's idea of an ethical focus seems to be keeping powerful models out of the hand of anyone unethical, which coincidentally is everyone except him.

tyre 6 hours ago

Yeah this would be a great point if it were true and they didn’t give Mythos access to companies to fix bugs, which they did and have.

It’s genuinely a difficult question. Not black and white. The models are really good at finding bugs, as demonstrated by people using Fable to reverse engineer. People make it sound like he’s just making it up.

throwaway63486 6 hours ago

hgoel 6 hours ago

codexon 5 hours ago

phoghed 6 hours ago

porphyra 6 hours ago

Dario's "ethical" look is also kinda sus. I hate to use ad hominem, but the dude's wife literally pitched a porn film to Epstein even after he was a convicted registered sex offender [1]. Dario is also really sinophobic (it is commonly claimed in Chinese AI circles that his former employment at Baidu triggered him so much that he harbors a personal grudge against the entire race).

[1] https://www.forbes.com/sites/alisondurkee/2026/08/14/who-is-...

tehlike 7 hours ago

Pretty much. That's even worse imho.

platinumrad 7 hours ago

Funny. Dario seems like the biggest snake in the industry to me and has leaned the hardest into doom marketing out of all of the influential leaders. With Altman (or Google), it's a transaction, and that's something I can live with.

tyre 6 hours ago

I just don’t see how people have looked at what has happened with Mythos and the deluge of fixes from companies, then come to this conclusion.

He has a really hard job. He errs on the side of conservatism in releasing and then people get Really Mad.

Safeguards on cybersecurity are not great for Anthropic revenue! As evidenced by people getting pissed, moving to Sol, and them having a smaller market for what Fable can do.

It’s clearly bad for revenue and not great advertising to say, “you can’t use this but here is a nerfed version that will annoy you and not solve important problems.”

adriand 6 hours ago

codexon 4 hours ago

SwellJoe 6 hours ago

drob518 7 hours ago

Gotta be honest that I’m tired of the “I hate Zuck and Meta so much” comments every time Meta does anything. Ditto Elon/X. Fine, I get it. I don’t like Zuck either. But the post is about Muse Spark 1.3. What do you think about that? If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.

duplessitous 7 hours ago

Technology doesn't just spring into being, there will always be comments on the organizations that developed it. If you don't like them or find them repetitive, it is far easier to collapse them and move on then bend a stranger to your will

canadaduane 7 hours ago

I get it, but the underlying problem is: we don't have a society-wide, effective solution to counterbalancing extractive systems. Lacking a reliable label, we have to constantly signal what's on the ingredients list.

drob518 7 hours ago

luckylion 7 hours ago

noduerme 7 hours ago

What I'm tired of is the top story (or five) on HN every day announcing Spark Opus Fable Grok Gemini v4.1i3-F. Like, who actually cares? Are people excited for the new benchmarks? Is it interesting to read the model cards? And look, part of my job is to use these things and part of my job is to pick EC2 servers, too. The front page of HN is increasingly resembling one of those endless AWS pricing lists.

And yeah, I don't like any of the people or companies building LLMs either. At least the griping is somewhat interesting by comparison. The model isn't news. The news on Hacker News is that other professionals feel the same way.

hadlock 6 hours ago

SyneRyder 6 hours ago

abjhn 5 hours ago

NamlchakKhandro 6 hours ago

monster_truck 7 hours ago

That's not how any of this works my man. Must be nice to think you live a life where neither has had a profound negative impact on your day to day

georgespencer 6 hours ago

> If you don’t like it […], then maybe just […] stay silent.

You might consider following your own advice.

dangoljames 7 hours ago

yeah, we should all just stfu because one internet dude is tired of hearing it

idiotsecant 7 hours ago

Not liking something because the embodiment of corporate malfeasance is a rational way to decide what products to support.

jesse_dot_id 7 hours ago

Staying silent is unfortunately how fascism festers.

whateveracct 7 hours ago

Zuck's bad PR is to blame here. Not the commenters. He should fix that.

Anduril makes this same complaint whenever their job posts get dumped on. Same idea. Fix your bad PR, buddies :)

troupo 7 hours ago

> But the post is about Muse Spark 1.3. What do you think about that?

That:

- like all models it was trained on stolen data

- additionally it was trained on Facebook users who were all opted in to AI training with a convoluted 10+ step process to opt-out of

> If you don’t like it because Meta made it, then maybe just don’t use it and stay silent.

Why should anyone stay silent?

badsectoracula 6 hours ago

> he seems to have the most ethical focus

He wants to build a tech-god kept in chains whose power he parcels out to the unwashed masses he deems worthy like some sort of high priest of intelligence.

And that is being charitable and going by the interpretation that he actually believes what he says.

im3w1l 6 hours ago

Well what do you want? Presenting clear, desirable, and achievable visions and trying to build consensus for how AI should develop is crucial at this point in time.

aftbit 7 hours ago

Okay but Muse Glimmer 30B is one of the best small open weight models today, and IMO the best from a US lab (only real comparison is Gemma4 dense right now).

tyre 6 hours ago

Totally fine with open weight, since other people can provide it and Meta isn’t making money. I’d use an AWS-hosted version.

HDBaseT 5 hours ago

Bluestein 7 hours ago

I am finding Poolside's a decent model.-

hadlock 5 hours ago

devy 7 hours ago

> I'm not an Amodei stan, but of all of these people he seems to have the most ethical focus. Again, not everything done perfectly and I have my gripes, but of the leaders of frontier labs, I'll vote with my money.

Amodei is NO Saint!!! He's the most savvy in drumming up the AI doomsday scenarios and haven't yet to apologized his failed forecast of Claude taking over 90% of the coding jobs.

samtheprogram 7 hours ago

He didn't say 90% of the coding jobs. He said LLMs would write 90% of the code. As in be LLM generated.

ls_stats 6 hours ago

Is it really hard to understand that there's no good guys? Amodei, Altman, Zuckerberg, Musk, etc. They all sound the same to me.

redox99 7 hours ago

Google, Zuck, Sama, Elon, Amodei (in no particular order).

They all suck. Pick your poison.

kenjackson 7 hours ago

They don't all suck equally.

Here's the order, from best to worst.

Amodei

Google

SamA

Zuck

Elon

redox99 7 hours ago

cactca 7 hours ago

scottyah 7 hours ago

porphyra 7 hours ago

runarberg 6 hours ago

anukin 5 hours ago

Anthropic is not exactly a saint either. I had a recent issue where they denied fable credits even though I was hospitalized during the claim period. I have annual plan with them. As much as everyone hates sama, I think OpenAI is much more of a company with good marketing and sales team.

optimalsolver 7 hours ago

If it was up to Dario we'd all be banned from using open-weight models, and we'd have to be investigated for PRC connections before sending our allotted five API queries a week.

jansport123 5 hours ago

no loyalty to any company - let them compete and then we get to choose.

ballon_monkey 5 hours ago

I'll happily pay for Grok, it's a great model. 4.6 often does better than Anthropic at coding and analysis where Anthropic fails for 'oh no cyber security, don't ask me to check if you're redacting passwords correctly in logs'. And it has no problem telling the truth where OpenAI / Anthropic don't want to upset the people on the left and will happily lie or avoid hard truths.

Edit: I get it. It's a hard pill to swallow. I understand people don't like Musk or Zuck. But it doesn't change the fact that you're being lied to and brainwashed.

dimgl 6 hours ago

I'll use both Muse Spark and Grok.

spiderfarmer 6 hours ago

Same with Grok.

loeg 7 hours ago

"Avoid generic tangents" / "Please don't complain about tangential annoyances."

_diyar 7 hours ago

How is this a tangential annoyance or a generic tangent?

> Meta announces they have a new model, demonstrating its capabilities.

> Parent comment states „regardless of this model‘s specific capabilities, if I can avoid it I will.“

loeg 7 hours ago

reaperducer 7 hours ago

"Avoid generic tangents" / "Please don't complain about tangential annoyances."

That's pretty much 90% of HN these days.

Apple releases a new iPhone? Here comes the flood of decade-old complaints about long-discontinued Mac butterfly keyboards and walled gardens.

Microsoft releases a new version of Windows? Here come the gripes about Azure.

Google changes something in GMail? Play Store!

It's like there's an army of bots out there determined to reduce the productivity of the Western tech bubble by diverting everyone into endless circular arguments about absolutely nothing of relevance to the topic at hand.

inferniac 5 hours ago

anthropic people have genuine delusions of grandeur, in a way they the worst of all the ai companies, definitely most cult-like

greatgib 5 hours ago

I hate Meta main business, but you have to admit that on the non business related and open source side, they have released amazing things that changed the world.

React for example.

And we could easily guess that there wouldn't have been so much open source models, and grand public experiments and free tools if llama models were not release to the general public.

bradlys 7 hours ago

What is the point of this comment?

jansport123 5 hours ago

someone expressing their view of meta which is on point by the way.

TacticalCoder 7 hours ago

Meta and Microsoft are two of the absolute worst evil companies on earth and Amodei is trying very hard to join them.

These Effective Altruists are despicable people: a bunch of thieves working to line up their own pockets while posturing as a force of good.

Remember that they schemed to not only present SBF as the 2nd coming of Christ (including in the NYT and in Forbes) but to also give him a voice after his scam had been uncovered. Thankfully, the judge didn't have any of this Effective Altruist bullshit.

SBF invested 500 millions of misappropriated funds in his buddy from the EA movement's Anthropic company (and, thankfully, the judge forced those shares to be sold: so SBF didn't get to be a billionaire).

You cannot hate enough people who say that harming others for the greater good is justified.

Then of course, already mentioned in this thread, there's the whole Epstein/Amodei's "I'm in the porn business" wife connection (where you don't need to squint much to see young women abused).

These kind of people are the absolute worst scum on this earth.

fouc 7 hours ago

How have they had a negative impact? How about google?

Yajirobe 7 hours ago

Enabled genocide in Myanmar

meerita 8 hours ago

I declined the use of cookies and everything went black. No content at all. Dissapointed.

improgrammer007 6 hours ago

All people here care about is hating Meta. Just look at the top voted comment. No one cares about the merits of the model, etc. HN has become nothing but an echo chamber.

dangoljames 7 hours ago

If it's from meta, pit h in the bin.

mgaunard 8 hours ago

They could have just called the article "struggling to remain relevant"

tonyhart7 8 hours ago

Meta is the last big tech come to AI race, so I would give prop to them for catching up