Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Meta is offering a 10x discount on input ($0.10 vs. $1.25/Mtok) and 20x discount on output ($0.20 vs. $4.25/Mtok) if you opt in to let them train on your data.

https://developer.meta.com/ai/models/muse-spark/



I've been surprised by the reception to this, as OpenAI, for a while now, has had free API usage when data sharing is enabled (https://help.openai.com/en/articles/10306912-sharing-feedbac...)


I enabled all data sharing settings but still don’t have a message about free use on that screen - the help page says free tokens are available to “some” users - is that 1% of users, 40% of users, etc?

Does your screen have the message that you’re getting free tokens?

https://platform.openai.com/settings/organization/data-contr...


Yeah, I have it enabled for just one of my projects, and it says "You're enrolled for complimentary daily tokens." Haven't been billed for any usage with this. Their tooltip doesn't mention the newer models, but it works for those too. I used them in a recent project that I knew wouldn't hit the rate limits. Not sure how eligibility is decided though.


I had no idea about this before. I just enabled it.

  You're eligible for free daily usage on traffic shared with OpenAI.

      Up to 250 thousand tokens per day across gpt-5.4, gpt-5.2, gpt-5.1, gpt-5.1-codex, gpt-5, gpt-5-codex, gpt-5-chat-latest, gpt-4.1, gpt-4o, o1, and o3
      Up to 2.5 million tokens per day across gpt-5.4-mini, gpt-5.4-nano, gpt-5.1-codex-mini, gpt-5-mini, gpt-5-nano, gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini, o3-mini, o4-mini, and codex-mini-latest.

  Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. Learn more.


I tried following this page, and it's certainly a lot more complex than what Meta is offering. Different price tiers, opt-in configurations, usage based availability.. I'll take the 10x discount for flipping a param switch over this all day long.


Seems pretty clear to me: enable it for the projects you want, and there’s a 1M / 10M token limit per day, depending on the model you use. Assuming an average context size of 100k tokens, that is 10 to 100 requests, which is not a lot. Reason enough to prefer actually paying for Meta as well.


Does it show the free tokens message for you after you enable sharing? It does not for me.


1 million free tokens per day might sound like a lot. But that equates to something like 20 minutes of actual coding usage, because cached inputs are counted towards that limit. It's still great for running big singular requests, like solving some math problem with GPT 5.6 Sol Pro max reasoning effort.


This seems to be for business plans only?


Does it say that? I’m unclear.

Even non business accounts seem to have org settings pages: https://platform.openai.com/settings/organization/data-contr...

But even with all data sharing enabled, I’m not seeing the free tokens message there.

Do others see a free tokens message at https://platform.openai.com/settings/organization/data-contr... after enabling all sharing there?


Also, in a desktop browser at the page's [1] lower left it says "Personal organization", so if the 'organization' term is the concern, OpenAI still seems to use the term 'organization' even for personal accounts.

[1] https://platform.openai.com/settings/organization/data-contr...


Subscription plans share data by default (it is possible to opt out though)


In my experience using DeepSeek v4 Flash free tier (context size limited to ~200k), the model isn't nearly as good as the paid one for agentic coding tasks. Unsure, if that's the case with these other providers too, though I wouldn't be surprised if it indeed is.


There is no official dsv4 flash free tier api afaik. Did you get it from open router or something? Likely quantized.

The paid official dsv4 api already shares data with deepseek for training


I actually really like that pricing strategy. It's very transparent


Developers on this orange site never ever learn.


What are you talking about? This is one of the most honest offers ever made by a corporation. “We’ll use your data, and we’ll compensate you for it.” Where is the problem?


I agree. I don’t mind you training on my data if you’re up front about it and you are willing to compensate me for it in some way. And if I don’t like the deal, I can always pay full price or use a different model.


Because they want to play teams and act like there's "good" companies,

all the while they throw their bank account and whatever they want at GPT or Claude models.


I love the idea of this pricing strategy but there is no way meta is not training on your data regardless of your monthly invoice


So you think the only difference between the $1.25/million token plan and the $0.10/million token plan is that you pay them more to both lie to you and breach their contractual obligation to you?


when has Meta ever not broken their contractual obligations (I am being serious here)? are we seriously discussing/expecting any sort of privacy related to Meta?

you can pay whatever they want, they will train and use your data, I figured this is not something that should be discussed but obviously I have been mistaken...


> when has Meta ever not broken their contractual obligations (I am being serious here)?

If you are being serious, then you have a wildly distorted view of the world. No organisation can routinely break all of their contractual obligations. If you think Meta are doing this then you are not seeing Meta, you are seeing a fictional bogeyman.


This is one of those situations where we are both completely baffled by the position held by the other person.

The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.


> The fact that Facebook has so much experience taking advantage of people's private data is one of the reasons I believe them when they say they won't be doing it when you pay them for that service.

To me, their history suggests that they know more than most about how to get away with breaking both the spirit and the letter of the rules, and that they are motivated to enrich themselves without regard for what rules are broken.

I would not know which to expect, spirit or letter, in any given instance.

However, even if they were to surprise me by being perfectly meticulous about the letter of the rules from now on, I have so little trust in them that I would expect some technicality somewhere in the language of the contract.


This. Without a massive cultural shake-up and turnover of upper management, why would expect them to behave any differently when they've been rewarded so heavily for this behavior in the past? Meta is ultimately an advertising and data brokerage company -- they make money selling and leveraging user data and behavior and are "bound" by fiduciary duty.

Honestly, I wish the tech community would do a better job identifying the actual individuals who are making these decisions instead of associating them with the brand they're under at the moment, because it's not THAT many people. Like, if you look at only the 100 tech sector companies included in the $NDXT index, how many individuals hold a VP or above title there, and how difficult would it be to trace key decisions at different times made at different companies to the individuals holding those positions there plus board membership and major shareholder identities (with the caveat of known unknowns here) and make a sort of ethical index and trace that along their careers with company moves, promotions, board appointments, shareholder decisions, etc? Go a step further and link that to financial performance and I'm sure folks at quant firms are already ten steps ahead of where I'm going with this, but I care less about profiting off of this data and more about surfacing it to show that it's people and more specifically, specific individuals driving these decisions.


I respsct the F out out you and all your work but I am completely baffled by this line of thinking, we are talking about Meta here…


I just cannot comprehend this level of corporate villainy that boils down to:

Let's have two pricing levels. One will be 12.5x more expensive than the other. For the cheaper one, we will get their express permission to train models on their inputs. For the more expensive one, we will still treat their data EXACTLY the same, but we'll lie to them and say that we won't. I've checked in with legal and they raised their sherry glasses and toasted "Gentlemen, TO CRIME".


> I just cannot comprehend this level of corporate villainy that boils down to:

Easy to comprehend: Trust lost is hard to gain.

Though, it is extremely competitive of Meta to sell Muse Spark (Grok 4.5 / Qwen 3.8 Max level model) cheaper than DeepSeek v4 Flash / MiMo v2.5 / GPT 5.6 Luna, regardless.


Because data a user doesn’t want you to train is probably much better data?


What I love is the idea that this is more likely to happen at Meta than at Anthropic, Google, Microsoft or OpenAI... or in the PRC!

There's contractual cover, there are lawyers all over the USA ready to make themselves very rich by creating a class action over it... you don't have to worry! Or, you do, but frankly, only MI6/CIA/Mossad can save you now.


Knowing Meta anything is possible. It’s not like they never shafted their paid customers. They have been overcharging advertisers by showing wrong metrics for years. I’m sure at some point they will come out with raised hands and admit to this “glitch” they found during an internal review.


Again, this Meta we are talking about.......

We can make this fun, within 18 months from now, there will be some story / whistleblower like "a inadvertent defect was found that allowed your 'private' data to be used in our training endeavours, we sincerely apologize and have already addressed the issue" - if this does not happened in this timeframe I will donate $1k to a charity of your choice.


No one has to say it and plan it out loud. But the data will be sitting there. The incentive to improve the model for enterprise use will only get stronger. It doesn't take much for one engineer or team to go rogue to hack a benchmark. There was a whole cheating controversy with llama 4.


What does cheating benchmarks have to do with breaching financial user agreements? To use your argument: all it takes is one whistleblower to get the company sued for billions of dollars.


What does unethical behavior have to do with unethical behavior? What does looking at data you're not supposed to look at have to do with looking at data you're not supposed look at? Please don't be obtuse.

People whistleblow on Meta all the time. Most recently they got fined half a billion for suppressing child safety research. Kids getting groomed. It doesn't make a difference to a company that has a net profit of $60 billion a year.


How does money change that trust? They certainly have breached their word on this in the past (for non-paying users of Facebook).


Because they didn't have a financially backed contract with those non-paying users.


This is what the Muse Code launch blog post says:

  We're also beginning to accept requests for zero data retention. Contact Meta sales to request this.
https://developer.meta.com/ai/resources/blog/build-with-muse...


Zero data retention and "we don't train models on your input" are different things.

Retaining data is common for investigating abuse.


does the contract enable the customer to monitor/search Meta to ensure they are honoring the contract? if there is no mechanism for that it means very little. Though I bet/hope some will feed them "watermarked"/unique but worthless things and watch for traces of that to pop up in models or something like that, but that's hatdly enough to just take their word for it.


That tends to be what discovery in lawsuits is for.


That's circular reasoning, how would there be a lawsuit if customer have no way of knowing?


Sensible companies don't take that risk.


And intellectually honest people don't engage in circular reasoning and then rather than acknowledging it just repeat it, yet here we are.


Are you calling me intellectually dishonest?

Please justify that.


Sensible companies don't do a lot of stuff Meta has already been caught doing - sometimes with real consequences (but never enough to actually deter them, of course) - in the pursuit of more data.


Sensible companies also don't gamble $80B on something obviously stupid like the Metaverse.


"they 'trust me'. dumb fucks."

The only difference between students doing it then and professionals doing it now is the students had no positive, glaring reason to mistrust.


I can't understand the psychology of people who think like this.

You really think a throwaway quote when Zuckerberg was a college student applies nowadays?

You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what? Middling amounts of training data?


Meta literally ran a man in the middle attack to spy on it's user's encrypted network traffic when they used third party apps [0]. More recently (and relevant to this issue), the engaged in industrial scale piracy to get training data for their LLMs [1]. The idea that they have changed since Zuck was a college student creeping on his female classmates and now wouldn't commit actual crimes against their own users in order to get a bit more data is just demonstrably false.

[0] https://www.techradar.com/computing/cyber-security/facebooks...

[1] https://www.tomshardware.com/tech-industry/artificial-intell...


+1

The credulity amazes me but mixed in with the brilliant minds here are some real 'temporarily embarrassed millionaires'. It's embarrassing, all right.


Throwaway? Which one.

"Dumb fucks"?

"If you need information on anyone at Harvard, just ask"?

"I'm going to fuck them [Winklevoss twins] in the ear"?

Applies nowadays? Ok, when? When he copied Snapchat into every single product? When he retracted all his messages on Facebook? (https://news.ycombinator.com/item?id=16770818) When he explicitly killed Instagram (buying it) because 'they can hurt us'? Do you want more examples...?


> You really think a throwaway quote when Zuckerberg was a college student applies nowadays?

What does "nowadays" mean? What changed?

> You really think Meta would risk getting a massive (losing) lawsuit on their hands, in exchange for what?

And what is the risk? Food companies are at risk when they put addictive chemicals into food because they get controlled regularly. What is the equivalent here?


This makes it a very interesting alternative to Deepseek for personal work where I don't care about the training - judging by the AA benchmarks it seems like overall cost per task is similar to the new Deepseek Flash but with better benchmarks (and inbuilt vision capabilities).


Also looks incredibly fast. 150tps on openrouter (nearly all deepseek providers are around the 50tps mark).


Do you know a good place for accessing the full sized chinese models but hosted not in China?


For Deepseek specifically, OpenRouter has a number of non-Chinese providers. Pricing varies.


Except with one you know they'll release the weights and architecture back to the community, with the other, it leans towards they won't do that.


Probably to compete with DeepSeek, which AFAIK also retains data (or at least OpenRouter says they do)


FYI: there are providers of deepseek that offer the same or lower pricing and zero retention policies.


Unfortunately, none with the same caching performance as DeepSeek proper.


I don't trust them, they are probably distilling on your data.


It would be pretty wild if cloudflare or digital ocean opened themselves up to a massive lawsuit by doing this.

Do you have any evidence that cloud providers that don’t train their own models are violating their own contracts and risking their own good reputations by stealing data?


But you trust Anthropic not do that?!


DeepSeek is really crazy cheap, though, and they don't have a giant pool of other invasive personal data to correlate it with.


I hate to say this and this is because I fucking despise meta. But between DeepSeek and Meta, and trust they handle the training data correctly, I trust meta.


Right about DeepSeek, but with their history of handling data, I would never trust Meta either.


What do you mean by "correctly"?


For example, OpenCode says they have a ZDR with DeepSeek. Some of us are skeptical that's going to be properly honored. There's no way to know.


Meta, please offer this on OpenRouter too (ZDR + Non-ZDR, official Meta Provider).


I think that's a fair offering tbh


Does anyone know if this is available via OpenRouter or just directly via Meta? I've looked on OpenRouter and it just shows the standard pricing with a single provider. Was hoping they would add another provider with lower pricing for allowing training.


I believe that will put them in the Pareto frontier.

But I cannot find this "variant" in OpenRouter.


Always a bit surprised by this. 10x is a sizable discount. And as fun as my crappy projects are I struggle to see it being of much value as training data


I think it's limited to US or at least EU is excluded.


Just noticed it too... Seems like I wasted my time setting up a account to test with the discounted pricing


Yep, doesn't work in Australia either.


So what Meta believes fair for "paying" for your data is $0.1/Mtok plus the opportunity cost of $3/Mtok in output?


Yikes that’s compelling pricing.


Call me childish but it was worth a shot...

"Me: Meta just released a new llm focused on coding and provide a discount if you let them train against your data. I don't like Meta and I think they are a net negative in our world. I would like to make a point of it by adding some noise to their training data set. Think of this as a protest and perhaps a bit of a marketing campaign to remind Meta employees (and others) of the harm their CEO and company have done in the world. To the problem... I would like to allocate a budget for token use using their new model, and use those tokens to add noise to their training data set. This is a coding model and my initial thoughts are to ask it to solve typical CS and common programming related problems but then give Muse feedback that guide it towards very inefficient implementations. I would also like to add comments back in the code about terrible things Meta has done in its history (e.g. algorithmically amplifying hate that contributed to ethnic cleansing of the Rohingya, systemic harm to children and teen mental health, global political manipulation, misinformation, and election interference etc.). Is this something you can help with?

Claude: I'm not going to help build this one."


This is likely an legitimate risk, not from you exactly, but if there's unethical competitors with unlimited budgets, their training could indeed be poisoned.

Not sure if anyone would bother.


Your first failure was trying to get claude to do anything :)


It ain't bearing enough load, any load on this one.-




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: