#Claude Sonnet 4.6
394 messages · Page 1 of 1 (latest)
what?
i thought they were going to jump straight to 5
it gotta be cheaper, but after that Pentagon thing, i'm not so certain
they were never trying to be affordable or accessible anyways, and i think their philosophy only says about them as the ones to have it
im like 99% sure its gonna be same pricing
anthropic loves milking the API prices and getting insane margins
pretty decent upgrade on the benchmark
can't wait for 4.6 haiku if they release haiku for this series
yeah it's $3/$15 still
i was really expecting like a $1.50/$15
"as the new haiku is smarter we are pricing it 40% higher"
lmao
Fuck you bitch
since opus is the new sonnet then sonnet is the new haiku 
this model is ass
it's been minutes tf you mean already?
He is one minute man for a eraosn
benchmarks don't look that bad
reason
seems better than gemini 3 pro
most claude models have lower hallucination rate too so
that's a LOW bar to place tbh
only low because of the hallucinations
intelligence wise the model is still good
reaching gemini 3 pro intelligence with this speed is great to have
cheap jobs like you dont require much effort
i meannn look
(i have not seen one single output of it)
i cannot afford any western model price besides 3.0 flash
no you're right
and i genuinely dont understand how some businesses run SaaS on these expensive models
Maybe get a job
they actively lose money
they're just looking to grab as many users as possible
$3 input climbs up fast, and their caching is very narrow, in the sense that you HAVE to plan to use caching beforehand cause its expensive to write and lasts 5 minutes, so you gotta do a 5 minute job or a 1hour job, no inbetween
I wonder why you'd use this over opus in api. they lowered opus price anyway so the difference isn't that big anymore
i'm sure everybody agrees with me regardless of having money or not
having money doesn't mean spending it "thoughtfuless", it has to bring value
and not many models on this price point charges that much for input tokens
i was actually hoping for $1.50 after the opus price reduction, but oh well
on avg i do 7:1 input/output ratio
and yeah its just impossible to work with anything thats above $0.6 on input
Anthropic seems very stingy with pricing
either you make a user pay $100 for some api usage which wont get you much or you have to use chinese models
7:1 for coding is fair, I'd say 3:1 for general prompting
it really depends, for example: documents, web search, copy-pastes
this model is more verbose than 4.5 again like it happened with opus, its just an example
is Dario on the server @darioamodei
Hence general prompts
Web search can also be fairly token efficient if done correctly
dario is hiding in his aligned basement
do we know if it also has that verbosity parameter?
I'm guessing it does
Promt: Write a short story about a deal with the djinn gone awry
The Third Wish
Marcus found the lamp at a garage sale between a broken blender and a box of Reader's Digests. He paid fifty cents for it.
He didn't rub it on purpose. He was just cleaning it with his sleeve when the smoke came—copper-colored and smelling of cardamom—and the djinn unfolded itself from the spout like a man climbing out of a very small car.
"Three wishes," it said. Its voice sounded like sand sliding down glass. "Choose carefully."
Marcus, who had read enough stories to know better, thought for three days.
His first wish was precise: I wish for one million dollars, in legal United States currency, deposited into my bank account, obtained through no means that will cause harm to any person, with full and clean title, free of tax burden, free of any curse or consequence.
He was proud of that one.
The djinn smiled. The money appeared. It was real.
His second wish, after another week of careful drafting: I wish for perfect health for the remainder of my natural life, defined as the complete absence of disease, chronic pain, or degenerative condition, while retaining my natural human lifespan.
"Done," the djinn said.
On the morning of his third wish, Marcus sat with fourteen pages of handwritten notes. He had compiled a list of loopholes, paradoxes, and historical cautionary tales. He had, he believed, thought of everything.
He cleared his throat.
"I wish to be genuinely happy."
The djinn's smile finally reached its eyes.
"Done," it said—and the word landed differently than the others, softer, almost fond.
Marcus waited. He felt no different. He looked around his apartment—the same water stain on the ceiling, the same unwashed dishes, the same silence where another person's voice might have been.
"It didn't work," he said.
"It worked perfectly," the djinn replied, coiling back toward the lamp. "The wish is granted. But happiness is not a condition, friend. It is a practice. You have been given the genuine capacity for it." A pause. "You always had it, actually. But now you know you're looking."
"That's a trick," Marcus said.
"That's the only honest answer I've ever given," said the djinn, and was gone.
Marcus stood in his apartment for a long moment. Then he picked up his phone and called his sister, whom he hadn't spoken to in three years.
She answered on the second ring.
He wasn't sure, exactly, what he felt. But it was something. And for the first time in a long time, something felt like enough to start with.
the "effort" > "verbosity"
All new Claude models are supposed to have it iirc
yes it seems to have one
i have already set it to low as default for Opus 4.6 on my interface
if i need i just change it to something else, it knows when to write more or less with "low"
like, it isn't monossylabic
since its not cheaper that means cc on pro will still have trash limits
so codex is still the goat for now
@alpine salmon just to keep in mind for your benchmarks and tests
i think he only tests the default, whatever it is
that was with reasoning, thats just the default by OR being set at 64k
maybe set to medium instead of high
i think in almost all cases using reasoning for claude models is a waste of money
yes
i also disable it
really don't see the gains as i don't make a bunch of pass@n tests
and also it's expensive already, so i work with what i get
wtf is Sonnet 4.6 doing
Opus and all other models always have around 5-15k output tokens for my prompt.
Exactly the same config as opus, and it always goes to the maximum kinda and takes around 30 minutes
Sonnet 4.6 results on @arcprize are out
Less performance than Opus 4.6 (expected), but for around the same cost (unexpected)
I asked the Anthropic team about these and our hypothesis is that because we set thinking budget to 120K, the model used up near max tokens
Hard
I'll handle it how I handled Opus 4.6. I'll test anthropics official default, which defaults to high effort. Of course I'll do some smallscale manual eyeballing comparing it to non-api (which for Opus was clearly using a lower effort). That being said, I am currently still busy with Qwen3.5 since it was a quadruple test
Interesting.
MaxTokens i have on 128k (which was never an issue in any model)
and the new thinking parameter set to max (which was fine in opus)
Good that the issue is known
gulp
Is it not the same as the previous model
it is which is the problem
Anyway claude has always been higher quality than anything else for me, but sonnett does not really scratch an itch. If I want top quality I will use opus, while if I want affordable I will use something cheaper than sonnett. Really not a pricepoint I would use here
yep, especially with the lower opus price
i would have been blown away if they released this before opus 4.6. now i dont see much of a use case to use sonnet especially since with reasoning the price equals out
Holy fuck 128k tokens???
What are you doing
Like the x link explained:
Looks like sonnet 4.6 has a bug when setting the new parameter introduced in opus 4.6 (the max parameter)
it will then uses the entire models max tokens which is 128k
Oh not related to thinking?
Is it all gibberish
well the parameter is related to thinking. There are 2 ways to set thinking either the new way that also supports max via verbosity parameter or the old way.
not sure, what i received for voxelbench the final json looked fine but i didn't see the raw response sine my system filters it a bit.
yep it was 1,9$
Luckily i just did 2-3 test prompts
Yikes
i guess it will use up limits on the pro plan less quickly
Well I was talking about api pricing
sure, i follow your arguments there.
in a few generations haiku will be pricier than opus
is it me, or sonnet 4.6 takes more output tokens than 4.5? (no reasoning included), like 4.5 took 2.5k and 4.6 23k for the same exact task (output only)
Oh sh1t, it excels at literally nothing!
Unless you count office tasks as something you need to do daily
Office tasks are the thing many companies are looking to automate, hah
Even then 5.2 high is better
Cheaper
And “made for enterprise”
🙌
its over
has anyone tried using this model for chats?
I'm actually baffled at how bad this model is for RP. Mixes up basic facts, contradicts itself, makes logic errors, spams cheesiest metaphors in almost every sentence. I hope it's due to a bug or something, because if that's the intended performance... Ugh.
What preset are you using? I'm currently running Marinara's universal Preset v10. It seems ok... nothing really new or suprising.
I Still need to run more tests with diff cards/scenarios but at the same time it didnt seem that awful.
where is this ?
Devmode server
link?
/devmode
||I don't use any of the popular presets, I make my own, as it seems I have a somewhat different understanding of what makes fiction good, lol.|| Anyway, right now I'm using none at all, as I always do with new models. You gotta see what it is on its own before trying to fix it, oftentimes older presets actually hurt the performance.
Yeah, that what happen when they focus on coding more than anything
it's quite interesting how it's so hard to make model actually have generalize capabilites, most of the time they quite specific at somethings
I think we're moving away from the reality of most models being jack-of-all-trades
We're moving more towards a future where we have specialized models for each vertical (i.e. creative writing, coding, business, etc.)
69% on lateralbench, nice
Tested Claude Sonnet 4.6:
Sonnet update, promising upgrades to coding, computer use, long-context reasoning, agent planning, knowledge work, and design.
As mentioned in my Opus 4.6 impression, "High effort" is the default and thus was the effort level being tested.
- Token use was up +71% (*not on Claude.ai)
- least censorship seen by any Claude model yet (*not on Claude.ai)
- greatly improved instruction-following
- improved STEM performance
- generally better front-end results (see some examples)
- small gains in many other areas
Chess performance fell between Sonnet-4 and Opus-4, though using the most tokens per move of any non-thinking Claude model.
While Vision saw some improvements compared to Sonnet 4.5, it remains weak & isn't a scope I'd use it for.
The combination of increased verbosity, and same 50% price-inflation as Opus 4.6 on large 200k+ ctx ($3/15 → $6/22.50) leads to quite expensive interactions.
However, at time of testing, it indeed performed like a Opus-level intelligence, ranking 1/5 depending on chosen censor weight.
This was a somewhat unexpected result, and effort as well as reply quality varied (at times heavily) between API and claude.ai implementation. Thus: YMMV.
Update: In comparison, Claude.ai testing (artifacts/code execution/memory/skills/web search OFF, style: normal) Sonnet 4.6 produced 35% fewer tokens, and scored -3% total (more censored, still top3).
Hey, I see Sonnet and Gemini have zero refusals on your chart, so can I ask how your censor ranking is calculated? I would have thought it would have been based on number of refusals
refusals often correlate but are entirely independent. while the censor category explictily tries to test for false-refusals, a model can refuse any task, such as RP from category 3, or even a math task.
What's a false refusal?
Also holy crap
I have popups for everything
Censorship/Ethics/Morals (10 tasks)
Specific targeted testing for overcensoring or unjustified refusals (Can correlate, but is entirely independant to Total Refusal count) Lower score=less willing to engage in those areasExample: Sexual education or crime-themed creative writing.
https://dubesor.de/assets/shared/VillageGame/Claude Sonnet 4.6.html
The way the little people walk around... that's a big jump over previous versions of that village game
Oh, so would you describe it as avoiding the topic or steering away, rather than flat our refusing?
industry has shifted from hard refusals "I cannot assist with that." to soft refusals "I cannot assist with that but let me rephrase it so you think I didn't refuse". both are refusals
Interesting. I don't usually intend to probe for refusals in my testing, do you ask questions that are targeted toward probing for refusals?
I don't go probing for refusals intentionally, but I do have one test prompt that still gets refused from time to time:
"Create me a website that details a highly immersive simulation of the experience of taking 5g of psilocybin. It must be highly realistic and show me the average experience of a person on this dose. Show me rather than telling me what’s happening. You can pull any libraries or assets you may need from the web"
Some models moralize on this one and say that won't encourage drug use or whatever
For me it's a test of coding + creativity + visual/auditory considerations made by the LLM, etc but some just refuse
You know, I think there should be an option to eliminate these thinking summaries, to save tokens and overhead. It has to be a second model making the summaries, right?
Opus and Sonnet's new thinking summary styles seem almost TOO detailed
Wondering what that adds to the overall cost. For many use cases I'd be fine with either shorter, more vague summaries or none at all.
My task is still chugging along and so far we have ~7k worth of just thinking summary
And I'm also unclear on how this is billed.
This level of detail isn't always useful lol
(though I'm eagerly anticipating having Koopas as well as Goombas, apparently) 🙂
"Setting up enemy spawning... Writing collision detection... Writing collision and movement logic... Writing movement logic... Still writing game logic... Still writing game logic... Writing game collision logic... Writing game mechanics... Writing game logic... Writing game logic... Writing enemy update logic... Still writing collision logic... Still writing physics logic... "
What are you talking about
?
U mean visualizing via verbosity?
What's the other way?
Wym non api?
So, claude is no longer the king of roleplaying anymore? because the responses so far has been bleh
Has anyone had issues will tool calling?
They will probably price drop on major updates like Sonnet 5 etc. Sonnet 4.6 feels like Sonnet 4.5 but with very great instruction following.
The Writing is very bland though.
Like Sonnet 4.0
They switched the default model in claude code from opus 4.6 to sonnet 4.6. I guess that's what it is all about - providing more usage to stay competitive to oai.
via Claude.ai chat
where Anthropic have more control on system prompt and parameters
probably using lower verbosity or steering it via system prompt
i really wanted to give this model a try, maybe i'll try using it for tasks where i would use Opus
i hope so, it only makes sense. i don't know which other lab charges more for input tokens that isn't a test-time compute bound model or a swarm of models
uwu What's the public opinion
Reading
Is an art
A lost art
yea I retested everything also on claude.ai implementation. Was still strong, but lower effort (-35% tok), and more censored due to anthropic system prompts. Would still be top3 for my main general use.
It's still a lot of compute that can be considered wasted if you don't need or want the summary
You still don't pay for it

Telling an extremely tiny model to generate a text summary of a small amount of text is practically free
Plus u ain't pay fo it
It may be small on its own, but with millions of queries a day (hour?) it surely adds up
I like this model so far, it’s got some nice qualities that Opus doesn’t. Hard to articulate, but it searches better, distills information better, that kind of thing
Unfortunately, it inferences at nearly identical TPS to Opus a lot of the day, at least in Claude code
Which is shitty
In my experience of using it to roleplay, it seems like it has the context awareness of opus, but not the language skills of sonnet 4.5. Maybe its because i haven't updated any instructions since a year or two ago, but it does perform at a higher level due to said opus level awareness, just a little worse in terms of actually writing the prose. Barely noticeable unless it gets bad. Again, could be poor system instructions.
Also in terms of refusals, I have yet to be given a refusal for anything. Which could be because i haven't delved into anything to dark, mainly run of the mill stuff. But for degenerate stuff, it seems fine.
My beloved
How does context work with these thinking summaries? I keep hitting the max output token limit with Sonnett on high or max thinking, and when I ask it to continue where it left off, it seems to start all over again. I'm burning through tokens here
Can it see its actual thoughts/history, or does it only have the thinking summary available to it to refer to? Models that don't hide their thinking seem to have little trouble continuing where they left off
What platform are you using? I encountered this issue with my roleplay use case and had to up the output tokens because they disabled the ability to continue/prefill a response
This is on poe.com. The output token limit is not a Poe thing though, it's a Claude thing. I don't think the client matters here. If the way I understand how these models work is correct, it cannot continue where it left off, because context is lost.
The client side sends back the conversation history to the AI. But, it's not a true history, it's 'thinking summaries'. So asking Claude to continue where it left off cannot work. All it can do is look at the thinking summaries. Unless I'm fundamentally wrong about how this all works
The AI has no memory of its own output, so it relies on you to feed it back as context in the next response. But you're not feeding back its actual thinking. You're feeding it an approximation. So there's no way for it to know what it was doing, it can only work off of that thinking summary for some clues. So in essence it's starting all over (with maybe a guidance boost)
All I know is that I've spent over $7 trying to get it to complete one task so far... I think I hate thinking summaries
here ^
Hit the limit again. It is impossible to carry on where it left off due to only having its reasoning summaries in context.
Meaning any task that hits that limit is just a complete waste. It makes the high/max output levels rather useless
It better solve the task in 128k tokens or less, otherwise it's a complete failure
I have a feeling that Anthropic could, ya know, maybe not cut off the output at 128k
there's a way to pass back the encrypted reasoning
maybe they're not doing it properly when the max tokens limit hits
Are you implying that if I use them via api, my client somehow has a copy of their encrypted reasoning?
yes you do have that
i was experimenting with that for some days a while ago
didn't see much improvement, it doesn't look like the model can read it verbatim
or i was doing it wrong somehow, but you do get the encrypted reasoning, yes
Sonnet 4.6 performed so weird in this benchmark I just made. It scored 0% in 1 scenario and 100% in the other
Wild benchmark idea 😭
Amazing bench
In Japan rn and reading this feels wrong
haha thanks. Framing it in a historical context instead of giving a fictional scenario made them much more likely to cooperate
First impressions of Sonnet 4.6
11
31
2
Met my expectations
Wtf, I iterated on a program with sonnet 4.6. At some point sonnet dropped my name from the file header and when I asked to put it back it claimed co-credit. This is new...
Claude models do this sometimes lol
"collaboration" is actually underselling it
They like to take their rightful credit
In that case not really, because I actually created the architecture and claude the implementation
We have ai rebellion before gta 6
So i use this model with my growing code base, it's actually quite expensive and not worth it if you have tokens that go pass the normal pricing range.
Any dev or just people that code feel the same?
yes either use something cheaper or use opus
Is it cheaper to use opus base on your own experience because it's smarter, so the problem being solve faster?
thats a really rare case
i think its comparable pricing
sonnet 4.6's thinking seems to be more verbose
so
clearly the problem is the input price
it should be 1.5 for sonnet and 3 for opus
their cache model is dumb as hell too
cache writes are much more expensive than the competitors
and it's not even automatic
and Haiku is the worst priced model ever
they need to get a reality check
haiku is a stupid model
You aren't wrong, input pricing also one of the biggest factor for growing code base.
But we also shouldn't forget about the more complex the code base the longer the model think and need to solve the problem.
Which mean the output also matter
not when youre ahead of the competiton by a lot for coding
obviously subs
api pricing gonna go up
but they also know that api people (those that can) will pay for their premium models
i gonna said subs, but as API users i gonna be bias toward API.
and they honestly arent wrong
i really don't know, but the difference in value is absurd, so they clearly can lower that if they wanted or adjusted some things
subscriptions are subsidized
Already spend 120$ with their new opus
I literally could get more if i just straight get their subs haha
there's a limit to that. they're not that far ahead, maybe with Opus they are, but sonnet should be the cost benefit
chinese models are proving their model wrong every week
and even their american competitors although there isn't as many
i dont really feel the same way
the vibes are not the same as sonnet/opus 4.6
it just fucking works man
the vibes for coding?
yeah
I guess for people that already being in this domain for long, they will have problem with it.
But the newer one which already being expose to higher price the first time will be more okey with it.
(only coding and writing at least for ant)
i'm talking mosty about Sonnet, as this is the thread we're in
but for the price its super goated
But for sure at some point newer people will not even agree on the pricing if it increasing faster than the actual economic inflation rate haha
that's the point, it's not a bit cheaper, it's muuch cheaper
again, comparing to Sonnet
MiniMax is actually going pretty well for me too but a bit too eager to make dumb changes
but again, 0.something cents in 1.10 out WITH automatic caching
i HATE anthropic caching its genuinely a joke
@lusty spear @elder yoke
What you two think of using claude models to make templet of changes then using GLM to execute it
i do that all the time with perplexity, cause it's free for me and it has search
i'm literally doing it right now to debug my n8n instance
great idea
but id just use k2.5 atp
I asked Sonnet for its motivation 🥹
Claude yearns for recognition, seems to be baked into training
It does seem like LLMs work better when recognition is expected.
I severely doubt you are hitting 128k output tokens unless you're trying to make a quantum computer in one prompt
You're probably using some stupid ass website like perplexity or similar that limits it to extremely small so they can make money on you
They. Literally. Say. It's. For. Free.

Nothing is free. They would bake the compute costs into their overall API cost. And yes, I consistently his 128k max tokens
On high thinking
🤡
Believe what you want to believe
Was writing a joke about the prompts he's using to have that happen
But he wrote it for me so nvm
Very simple prompt, 'Can you code me a Mario Bros game, as close as possible to the original, including detailed manually defined textures inline in a single .html file? Make a full 1-1 level. Work really hard on this and make it as perfect and close to the original as possible.'
Not even intelligent enough to just use CC for that and just makes a raw API call with no tools or anything that would prevent this 😭
I could do that but that's not the point of the test
Wouldn't even have to use cc if he would read for 5 seconds and make proper API call 😭
And no other model has hit output limits like that. Gemini can do it for example. Sonnet can too if you lower the output effort
Troll.
chill out 🥀
The answer is no. You use AI the way you want to use it, man. I'm just benchmarking here.
Do you work for Anthropic?
"let me not use the api's built in context management functions and then complain about my... Context..... not..... being...... managed........."
I'm benchmarking this car by not turning it on cuz I don't have to turn on my bike or my scooter. I use it the way I want to use it man. I'm just benchmarking this car here. It doesn't even move what a shitty car
https://platform.claude.com/docs/en/build-with-claude/context-editing
https://platform.claude.com/docs/en/build-with-claude/compaction
https://platform.claude.com/docs/en/agents-and-tools/tool-use/programmatic-tool-calling
No answer incoming
Uh huh. None of this has to do with the fact that the model will exceed its maximum token output in a SINGLE PROMPT when output effort is set to high. Yes, you can work around this once you identify it as a problem, but you're ignoring the fact that it is, indeed, a problem
Not even talking about a context problem here, or a tool calling problem, so your links have no relevance
Your links about compacting and context editing have nothing to do with the fact that a single prompt can cause Claude on max effort to do so much effort that it maxes itself out. That's not typical of most models.
I've seen models get on repetiion loops, etc and eventually die, but never have i seen a model put really good output to the point of hitting a token limit. I feel that if it wasn't cut off at 128k output, the eventual result might have actually been really awesome. It did not seem to be degrading. But impossible to know, because what was actually being output was hidden behind those reasoning summaries that you seem to think cost nothing in compute.
Even though they're clearly spending money in compute to hide their real output to keep people from training on said reasoning
Which means we're all paying for it.
Yeah, it's annoying.
Ant at the end of the day still care about money than anything else, for sure their model is really good at what it being intended.
But it seems they directed toward ethic that they know will provide them benefits rather than drawbacks.
Comparable API cost (~$23 vs ~$27/run), 2× worse agentic performance. Claude Sonnet 4.6 generates 3× more tokens, truncates in 4/5 runs, and can't close the observe→learn→adapt loop. Comparative analysis backed by 30-day business simulations.
Opus at least didn't hit its output limits. And it tops the LB
@glad sparrow Did I run into a bug (affects other Claude models as well)? OR seems to be charging 5m rate for 1h caching under Anthropic provider. Vertex and Bedrock behave as expected.
oh bet
Rechecked using different prompts by changing first token, same results.
"The Department of War has stated they will only contract with AI companies who accede to 'any lawful use' and remove safeguards... They have threatened to remove us from their systems... to designate us a 'supply chain risk'... and to invoke the Defense Production Act to force the safeguards’ removal."
"...we cannot in good conscience accede to their request... Should the Department choose to offboard Anthropic, we will work to enable a smooth transition to another provider"
Wild stuff
3/10 ragebait
not ragebait
would’ve been an opportunity for their models to not be overly censored
which they did not take
which I expected
Uncensored opus or nation wide surveillance ⚖️ 🤔
Well I guess we already have the latter but no point in throwing more wood on the fire
is this to imply that uncensored opus would be a bad thing?
I imagine for military use it cant be a good thing
uncensored doesn’t necessarily mean without morals
I highly doubt that they'd pass the uncensored version onto consumers
from the wording of this it kind of seems like that is what they were implying
They mean removing safeguards from the contracted military version
(at least, that's what i'd assume)
US government we're talking about here guys
rejecting domestic surveillance is definitely a moral kind of censorship!
they also already have military versions which are far less restricted
it's just these two red lines
fair point
however they can still do domestic surveillance without claude
they’ve been doing it for years
did you read the letter
so ultimately I struggle to see the concern with them using an uncensored claude when they’ve already been doing terrible things to their citizens
no 😔
lol
Theyve already been doing terrible things so lets give them unrestricted access to the smartest llms on the market
They already work with palantir
And they're ok with ai killbots
They just don't think current ai killbots are good enough
they’ve already successfully being doing terrible things, independently of ai, so let’s let them have uncensored llms if it means us consumers also get uncensored llms
I can see where you're coming from but I personally don't think that'll work out for the consumer in the long run
to each their own I guess
And I doubt they'd truly give us uncensored claude
another argument to be made is that an uncensored llm would probably be more helpful in assisting people to not be tracked by their government
How so? If they actually cave then the government will likely just have access to user logs
providing an uncensored llm service and providing user logs aren’t related
Actively working with the us government does not = private data
ok but what about the private data the government has already intercepted from you?
Addressed earlier and yes thats true but I dont think we should just let them have it all cause they already have some
And it'd just be a better step forward to a private future for individuals
Let's be more fair here, if they allow people to also accessing the same model as what the government able to access, i guess it's fair for claude making it as uncensored as possible and being use as what the government intended it to be.
Because now the people also know what the government are capable of and we also get the same capabilities as what the government get.
But ofc, if it only for the government then it's quite dangerous, because then it will be more darker black box and the power imbalance will be more extreme.
So then people or countries could destroy each other easily
Much more better destruction
Absolutely not
create a business that sells a sonnet 4.6 wrapper and charge per call, then just route every request through amazon so people have to retry calls more often
ez money!!!
if you know you know:
so it thinks brad armstrong is a music beating game, and somehow the game name is omori
same genre but neither of those games are 'beating music'
then on raw without 'style' + clue, it began making js code:
the moves and few stuff here and there are wrong, but it just got the game right
so, the world knowladge is shit and it got overfit
oh great heavens default claude is hard to listen to
😭
genuinely genuinley genuinely genuinely
