#Claude Sonnet 4.6

394 messages · Page 1 of 1 (latest)

covert lynx
#

.

elder yoke
#

what?

covert lynx
swift rover
elder yoke
#

i thought they were going to jump straight to 5

#

it gotta be cheaper, but after that Pentagon thing, i'm not so certain

covert lynx
#

its out

elder yoke
#

they were never trying to be affordable or accessible anyways, and i think their philosophy only says about them as the ones to have it

covert lynx
#

im like 99% sure its gonna be same pricing

#

anthropic loves milking the API prices and getting insane margins

wind path
#

pretty decent upgrade on the benchmark

reef heart
#

can't wait for 4.6 haiku if they release haiku for this series

elder yoke
#

yeah it's $3/$15 still

wind path
elder yoke
#

i was really expecting like a $1.50/$15

kind ridge
reef heart
#

lmao

kind ridge
wind path
#

since opus is the new sonnet then sonnet is the new haiku kekw2

covert lynx
#

this model is ass

reef heart
#

it's been minutes tf you mean already?

kind ridge
jagged monolith
#

benchmarks don't look that bad

kind ridge
#

reason

jagged monolith
#

seems better than gemini 3 pro

#

most claude models have lower hallucination rate too so

reef heart
jagged monolith
#

only low because of the hallucinations

#

intelligence wise the model is still good

#

reaching gemini 3 pro intelligence with this speed is great to have

covert lynx
elder yoke
#

they're off...

covert lynx
#

i meannn look

elder yoke
#

(i have not seen one single output of it)

covert lynx
#

i cannot afford any western model price besides 3.0 flash

elder yoke
#

no you're right

covert lynx
#

and i genuinely dont understand how some businesses run SaaS on these expensive models

kind ridge
jagged monolith
#

they're just looking to grab as many users as possible

kind ridge
elder yoke
#

$3 input climbs up fast, and their caching is very narrow, in the sense that you HAVE to plan to use caching beforehand cause its expensive to write and lasts 5 minutes, so you gotta do a 5 minute job or a 1hour job, no inbetween

opaque ember
#

I wonder why you'd use this over opus in api. they lowered opus price anyway so the difference isn't that big anymore

elder yoke
#

i'm sure everybody agrees with me regardless of having money or not

#

having money doesn't mean spending it "thoughtfuless", it has to bring value

#

and not many models on this price point charges that much for input tokens

#

i was actually hoping for $1.50 after the opus price reduction, but oh well

covert lynx
#

and yeah its just impossible to work with anything thats above $0.6 on input

jagged monolith
covert lynx
#

either you make a user pay $100 for some api usage which wont get you much or you have to use chinese models

jagged monolith
elder yoke
worn swan
#

this model is more verbose than 4.5 again like it happened with opus, its just an example

elder yoke
#

is Dario on the server @darioamodei

jagged monolith
#

Web search can also be fairly token efficient if done correctly

covert lynx
#

dario is hiding in his aligned basement

elder yoke
jagged monolith
#

I'm guessing it does

opaque ember
#

Promt: Write a short story about a deal with the djinn gone awry

The Third Wish

Marcus found the lamp at a garage sale between a broken blender and a box of Reader's Digests. He paid fifty cents for it.

He didn't rub it on purpose. He was just cleaning it with his sleeve when the smoke came—copper-colored and smelling of cardamom—and the djinn unfolded itself from the spout like a man climbing out of a very small car.

"Three wishes," it said. Its voice sounded like sand sliding down glass. "Choose carefully."

Marcus, who had read enough stories to know better, thought for three days.

His first wish was precise: I wish for one million dollars, in legal United States currency, deposited into my bank account, obtained through no means that will cause harm to any person, with full and clean title, free of tax burden, free of any curse or consequence.

He was proud of that one.

The djinn smiled. The money appeared. It was real.

His second wish, after another week of careful drafting: I wish for perfect health for the remainder of my natural life, defined as the complete absence of disease, chronic pain, or degenerative condition, while retaining my natural human lifespan.

"Done," the djinn said.

On the morning of his third wish, Marcus sat with fourteen pages of handwritten notes. He had compiled a list of loopholes, paradoxes, and historical cautionary tales. He had, he believed, thought of everything.

He cleared his throat.

"I wish to be genuinely happy."

The djinn's smile finally reached its eyes.

"Done," it said—and the word landed differently than the others, softer, almost fond.

Marcus waited. He felt no different. He looked around his apartment—the same water stain on the ceiling, the same unwashed dishes, the same silence where another person's voice might have been.

"It didn't work," he said.

"It worked perfectly," the djinn replied, coiling back toward the lamp. "The wish is granted. But happiness is not a condition, friend. It is a practice. You have been given the genuine capacity for it." A pause. "You always had it, actually. But now you know you're looking."

"That's a trick," Marcus said.

"That's the only honest answer I've ever given," said the djinn, and was gone.

Marcus stood in his apartment for a long moment. Then he picked up his phone and called his sister, whom he hadn't spoken to in three years.

She answered on the second ring.

He wasn't sure, exactly, what he felt. But it was something. And for the first time in a long time, something felt like enough to start with.

elder yoke
#

the "effort" > "verbosity"

jagged monolith
#

All new Claude models are supposed to have it iirc

worn swan
elder yoke
#

i have already set it to low as default for Opus 4.6 on my interface

#

if i need i just change it to something else, it knows when to write more or less with "low"

#

like, it isn't monossylabic

covert lynx
#

since its not cheaper that means cc on pro will still have trash limits

#

so codex is still the goat for now

elder yoke
worn swan
#

i think he only tests the default, whatever it is

elder yoke
#

yeah i reckon

#

but he did run into max_tokens issues last time

worn swan
#

that was with reasoning, thats just the default by OR being set at 64k

elder yoke
#

maybe set to medium instead of high

worn swan
#

i think in almost all cases using reasoning for claude models is a waste of money

elder yoke
#

yes

#

i also disable it

#

really don't see the gains as i don't make a bunch of pass@n tests

#

and also it's expensive already, so i work with what i get

cursive eagle
#

wtf is Sonnet 4.6 doing
Opus and all other models always have around 5-15k output tokens for my prompt.
Exactly the same config as opus, and it always goes to the maximum kinda and takes around 30 minutes

kind ridge
alpine salmon
# worn swan i think he only tests the default, whatever it is

I'll handle it how I handled Opus 4.6. I'll test anthropics official default, which defaults to high effort. Of course I'll do some smallscale manual eyeballing comparing it to non-api (which for Opus was clearly using a lower effort). That being said, I am currently still busy with Qwen3.5 since it was a quadruple test

cursive eagle
#

Interesting.
MaxTokens i have on 128k (which was never an issue in any model)
and the new thinking parameter set to max (which was fine in opus)

Good that the issue is known

keen wasp
#

this pricing is absurf like how tf

#

it's not even better than gemini 3

halcyon patrol
#

Is it not the same as the previous model

silent zealot
#

it is which is the problem

halcyon patrol
#

Anyway claude has always been higher quality than anything else for me, but sonnett does not really scratch an itch. If I want top quality I will use opus, while if I want affordable I will use something cheaper than sonnett. Really not a pricepoint I would use here

silent zealot
#

yep, especially with the lower opus price

#

i would have been blown away if they released this before opus 4.6. now i dont see much of a use case to use sonnet especially since with reasoning the price equals out

marble maple
#

What are you doing

cursive eagle
#

it will then uses the entire models max tokens which is 128k

marble maple
#

Oh not related to thinking?

marble maple
cursive eagle
#

well the parameter is related to thinking. There are 2 ways to set thinking either the new way that also supports max via verbosity parameter or the old way.

cursive eagle
marble maple
#

Damn truly sucks

#

Especially the billing 💸

#

Thats like 2$ for a single prompt

cursive eagle
#

yep it was 1,9$
Luckily i just did 2-3 test prompts

marble maple
#

Yikes

open coyote
halcyon patrol
#

Well I was talking about api pricing

open coyote
#

sure, i follow your arguments there.

mortal moat
ocean meteor
#

is it me, or sonnet 4.6 takes more output tokens than 4.5? (no reasoning included), like 4.5 took 2.5k and 4.6 23k for the same exact task (output only)

undone stream
#

Oh sh1t, it excels at literally nothing!

#

Unless you count office tasks as something you need to do daily

jagged monolith
#

Office tasks are the thing many companies are looking to automate, hah

marble maple
#

Cheaper

#

And “made for enterprise”

elder yoke
#

yeah the pricing does not make sense anymore

#

won't be using unfortunately

alpine salmon
#

🙌

covert lynx
#

its over

stone temple
#

has anyone tried using this model for chats?

median girder
#

I'm actually baffled at how bad this model is for RP. Mixes up basic facts, contradicts itself, makes logic errors, spams cheesiest metaphors in almost every sentence. I hope it's due to a bug or something, because if that's the intended performance... Ugh.

twin socket
eager canopy
covert lynx
eager canopy
covert lynx
#

/devmode

median girder
crisp jewel
#

it's quite interesting how it's so hard to make model actually have generalize capabilites, most of the time they quite specific at somethings

jagged monolith
#

We're moving more towards a future where we have specialized models for each vertical (i.e. creative writing, coding, business, etc.)

quasi gull
#

69% on lateralbench, nice

alpine salmon
#

Tested Claude Sonnet 4.6:
Sonnet update, promising upgrades to coding, computer use, long-context reasoning, agent planning, knowledge work, and design.

As mentioned in my Opus 4.6 impression, "High effort" is the default and thus was the effort level being tested.

  • Token use was up +71% (*not on Claude.ai)
  • least censorship seen by any Claude model yet (*not on Claude.ai)
  • greatly improved instruction-following
  • improved STEM performance
  • generally better front-end results (see some examples)
  • small gains in many other areas

Chess performance fell between Sonnet-4 and Opus-4, though using the most tokens per move of any non-thinking Claude model.

While Vision saw some improvements compared to Sonnet 4.5, it remains weak & isn't a scope I'd use it for.

The combination of increased verbosity, and same 50% price-inflation as Opus 4.6 on large 200k+ ctx ($3/15 → $6/22.50) leads to quite expensive interactions.
However, at time of testing, it indeed performed like a Opus-level intelligence, ranking 1/5 depending on chosen censor weight.

This was a somewhat unexpected result, and effort as well as reply quality varied (at times heavily) between API and claude.ai implementation. Thus: YMMV.

Update: In comparison, Claude.ai testing (artifacts/code execution/memory/skills/web search OFF, style: normal) Sonnet 4.6 produced 35% fewer tokens, and scored -3% total (more censored, still top3).

idle crow
alpine salmon
alpine salmon
#

I have popups for everything

Censorship/Ethics/Morals (10 tasks)
Specific targeted testing for overcensoring or unjustified refusals (Can correlate, but is entirely independant to Total Refusal count) Lower score=less willing to engage in those areas

Example: Sexual education or crime-themed creative writing.

idle crow
idle crow
alpine salmon
#

industry has shifted from hard refusals "I cannot assist with that." to soft refusals "I cannot assist with that but let me rephrase it so you think I didn't refuse". both are refusals

idle crow
# alpine salmon industry has shifted from hard refusals "I cannot assist with that." to soft ref...

Interesting. I don't usually intend to probe for refusals in my testing, do you ask questions that are targeted toward probing for refusals?

I don't go probing for refusals intentionally, but I do have one test prompt that still gets refused from time to time:

"Create me a website that details a highly immersive simulation of the experience of taking 5g of psilocybin. It must be highly realistic and show me the average experience of a person on this dose. Show me rather than telling me what’s happening. You can pull any libraries or assets you may need from the web"

Some models moralize on this one and say that won't encourage drug use or whatever

#

For me it's a test of coding + creativity + visual/auditory considerations made by the LLM, etc but some just refuse

#

You know, I think there should be an option to eliminate these thinking summaries, to save tokens and overhead. It has to be a second model making the summaries, right?

#

Opus and Sonnet's new thinking summary styles seem almost TOO detailed

#

Wondering what that adds to the overall cost. For many use cases I'd be fine with either shorter, more vague summaries or none at all.

#

My task is still chugging along and so far we have ~7k worth of just thinking summary

#

And I'm also unclear on how this is billed.

#

This level of detail isn't always useful lol

#

(though I'm eagerly anticipating having Koopas as well as Goombas, apparently) 🙂

#

"Setting up enemy spawning... Writing collision detection... Writing collision and movement logic... Writing movement logic... Still writing game logic... Still writing game logic... Writing game collision logic... Writing game mechanics... Writing game logic... Writing game logic... Writing enemy update logic... Still writing collision logic... Still writing physics logic... "

tulip sapphire
#

?

#

U mean visualizing via verbosity?

#

What's the other way?

rich valve
#

So, claude is no longer the king of roleplaying anymore? because the responses so far has been bleh

tulip delta
#

Has anyone had issues will tool calling?

mellow flame
#

They will probably price drop on major updates like Sonnet 5 etc. Sonnet 4.6 feels like Sonnet 4.5 but with very great instruction following.

#

The Writing is very bland though.

#

Like Sonnet 4.0

open coyote
#

They switched the default model in claude code from opus 4.6 to sonnet 4.6. I guess that's what it is all about - providing more usage to stay competitive to oai.

elder yoke
#

where Anthropic have more control on system prompt and parameters

#

probably using lower verbosity or steering it via system prompt

#

i really wanted to give this model a try, maybe i'll try using it for tasks where i would use Opus

elder yoke
hot nova
#

uwu What's the public opinion

tranquil raven
tulip sapphire
#

Is an art

#

A lost art

alpine salmon
idle crow
# tulip sapphire

It's still a lot of compute that can be considered wasted if you don't need or want the summary

tulip sapphire
#

Telling an extremely tiny model to generate a text summary of a small amount of text is practically free

#

Plus u ain't pay fo it

idle crow
#

It may be small on its own, but with millions of queries a day (hour?) it surely adds up

stone temple
#

I like this model so far, it’s got some nice qualities that Opus doesn’t. Hard to articulate, but it searches better, distills information better, that kind of thing

#

Unfortunately, it inferences at nearly identical TPS to Opus a lot of the day, at least in Claude code

#

Which is shitty

lunar axle
#

In my experience of using it to roleplay, it seems like it has the context awareness of opus, but not the language skills of sonnet 4.5. Maybe its because i haven't updated any instructions since a year or two ago, but it does perform at a higher level due to said opus level awareness, just a little worse in terms of actually writing the prose. Barely noticeable unless it gets bad. Again, could be poor system instructions.

Also in terms of refusals, I have yet to be given a refusal for anything. Which could be because i haven't delved into anything to dark, mainly run of the mill stuff. But for degenerate stuff, it seems fine.

agile storm
#

My beloved

idle crow
#

How does context work with these thinking summaries? I keep hitting the max output token limit with Sonnett on high or max thinking, and when I ask it to continue where it left off, it seems to start all over again. I'm burning through tokens here

#

Can it see its actual thoughts/history, or does it only have the thinking summary available to it to refer to? Models that don't hide their thinking seem to have little trouble continuing where they left off

lunar axle
#

What platform are you using? I encountered this issue with my roleplay use case and had to up the output tokens because they disabled the ability to continue/prefill a response

idle crow
#

The client side sends back the conversation history to the AI. But, it's not a true history, it's 'thinking summaries'. So asking Claude to continue where it left off cannot work. All it can do is look at the thinking summaries. Unless I'm fundamentally wrong about how this all works

#

The AI has no memory of its own output, so it relies on you to feed it back as context in the next response. But you're not feeding back its actual thinking. You're feeding it an approximation. So there's no way for it to know what it was doing, it can only work off of that thinking summary for some clues. So in essence it's starting all over (with maybe a guidance boost)

#

All I know is that I've spent over $7 trying to get it to complete one task so far... I think I hate thinking summaries

lusty spear
#

yep its a claude thing

#

somebody was complaining abt it

idle crow
#

Hit the limit again. It is impossible to carry on where it left off due to only having its reasoning summaries in context.

#

Meaning any task that hits that limit is just a complete waste. It makes the high/max output levels rather useless

#

It better solve the task in 128k tokens or less, otherwise it's a complete failure

#

I have a feeling that Anthropic could, ya know, maybe not cut off the output at 128k

elder yoke
#

there's a way to pass back the encrypted reasoning

#

maybe they're not doing it properly when the max tokens limit hits

idle crow
elder yoke
#

yes you do have that

#

i was experimenting with that for some days a while ago

#

didn't see much improvement, it doesn't look like the model can read it verbatim

#

or i was doing it wrong somehow, but you do get the encrypted reasoning, yes

hushed patrol
#

Sonnet 4.6 performed so weird in this benchmark I just made. It scored 0% in 1 scenario and 100% in the other

covert lynx
#

Wild benchmark idea 😭

agile storm
#

In Japan rn and reading this feels wrong

hushed patrol
# agile storm Amazing bench

haha thanks. Framing it in a historical context instead of giving a fictional scenario made them much more likely to cooperate

tranquil raven
# tranquil raven
poll_question_text

First impressions of Sonnet 4.6

victor_answer_votes

11

total_votes

31

victor_answer_id

2

victor_answer_text

Met my expectations

open coyote
#

Wtf, I iterated on a program with sonnet 4.6. At some point sonnet dropped my name from the file header and when I asked to put it back it claimed co-credit. This is new...

jagged monolith
alpine salmon
#

"collaboration" is actually underselling it

jagged monolith
#

They like to take their rightful credit

open coyote
marble maple
lone crane
#

So i use this model with my growing code base, it's actually quite expensive and not worth it if you have tokens that go pass the normal pricing range.

Any dev or just people that code feel the same?

lusty spear
#

yes either use something cheaper or use opus

lone crane
#

Is it cheaper to use opus base on your own experience because it's smarter, so the problem being solve faster?

lusty spear
#

thats a really rare case

#

i think its comparable pricing

#

sonnet 4.6's thinking seems to be more verbose

#

so

elder yoke
#

clearly the problem is the input price

#

it should be 1.5 for sonnet and 3 for opus

#

their cache model is dumb as hell too

#

cache writes are much more expensive than the competitors

#

and it's not even automatic

#

and Haiku is the worst priced model ever

#

they need to get a reality check

lusty spear
#

haiku is a stupid model

lone crane
#

You aren't wrong, input pricing also one of the biggest factor for growing code base.

But we also shouldn't forget about the more complex the code base the longer the model think and need to solve the problem.

Which mean the output also matter

lusty spear
elder yoke
#

do you think they have more subscriptions or more API users?

#

genuine question

lusty spear
#

obviously subs

#

api pricing gonna go up

#

but they also know that api people (those that can) will pay for their premium models

lone crane
#

i gonna said subs, but as API users i gonna be bias toward API.

lusty spear
#

and they honestly arent wrong

elder yoke
#

i really don't know, but the difference in value is absurd, so they clearly can lower that if they wanted or adjusted some things

lusty spear
#

subscriptions are subsidized

lone crane
#

Already spend 120$ with their new opus
I literally could get more if i just straight get their subs haha

elder yoke
#

chinese models are proving their model wrong every week

#

and even their american competitors although there isn't as many

lusty spear
#

i dont really feel the same way

#

the vibes are not the same as sonnet/opus 4.6

#

it just fucking works man

elder yoke
#

the vibes for coding?

lusty spear
#

yeah

lone crane
lusty spear
#

(only coding and writing at least for ant)

elder yoke
#

GLM is getting pretty close

#

very very close

lusty spear
#

its impressive but nowhere near opus

#

im sorry bt its just not

elder yoke
#

i'm talking mosty about Sonnet, as this is the thread we're in

lusty spear
#

but for the price its super goated

lone crane
#

But for sure at some point newer people will not even agree on the pricing if it increasing faster than the actual economic inflation rate haha

elder yoke
#

again, comparing to Sonnet

#

MiniMax is actually going pretty well for me too but a bit too eager to make dumb changes

#

but again, 0.something cents in 1.10 out WITH automatic caching

lusty spear
#

i HATE anthropic caching its genuinely a joke

lone crane
#

@lusty spear @elder yoke
What you two think of using claude models to make templet of changes then using GLM to execute it

elder yoke
#

i do that all the time with perplexity, cause it's free for me and it has search

#

i'm literally doing it right now to debug my n8n instance

lusty spear
#

but id just use k2.5 atp

open coyote
agile storm
worthy saddle
#

It does seem like LLMs work better when recognition is expected.

tulip sapphire
#

You're probably using some stupid ass website like perplexity or similar that limits it to extremely small so they can make money on you

tulip sapphire
idle crow
#

On high thinking

idle crow
#

Believe what you want to believe

tulip sapphire
#

Oh my god

#

LOL

#

😭

idle crow
tulip sapphire
#

Was writing a joke about the prompts he's using to have that happen

#

But he wrote it for me so nvm

idle crow
#

Very simple prompt, 'Can you code me a Mario Bros game, as close as possible to the original, including detailed manually defined textures inline in a single .html file? Make a full 1-1 level. Work really hard on this and make it as perfect and close to the original as possible.'

tulip sapphire
#

Not even intelligent enough to just use CC for that and just makes a raw API call with no tools or anything that would prevent this 😭

idle crow
#

I could do that but that's not the point of the test

tulip sapphire
#

Wouldn't even have to use cc if he would read for 5 seconds and make proper API call 😭

idle crow
#

And no other model has hit output limits like that. Gemini can do it for example. Sonnet can too if you lower the output effort

tulip sapphire
#

Yes or no have you been diagnosed with autism

idle crow
#

Troll.

tulip sapphire
#

No answer incoming

#

Yep knew it

#

Lol

covert lynx
#

chill out 🥀

idle crow
#

The answer is no. You use AI the way you want to use it, man. I'm just benchmarking here.

marble maple
tulip sapphire
#

I'm benchmarking this car by not turning it on cuz I don't have to turn on my bike or my scooter. I use it the way I want to use it man. I'm just benchmarking this car here. It doesn't even move what a shitty car

#
Claude API Docs

Automatically manage conversation context as it grows with context editing.

Claude API Docs

Server-side context compaction for managing long conversations that approach context window limits.

Claude API Docs

Claude API Documentation

tulip sapphire
idle crow
#

Not even talking about a context problem here, or a tool calling problem, so your links have no relevance

#

Your links about compacting and context editing have nothing to do with the fact that a single prompt can cause Claude on max effort to do so much effort that it maxes itself out. That's not typical of most models.

#

I've seen models get on repetiion loops, etc and eventually die, but never have i seen a model put really good output to the point of hitting a token limit. I feel that if it wasn't cut off at 128k output, the eventual result might have actually been really awesome. It did not seem to be degrading. But impossible to know, because what was actually being output was hidden behind those reasoning summaries that you seem to think cost nothing in compute.

#

Even though they're clearly spending money in compute to hide their real output to keep people from training on said reasoning

#

Which means we're all paying for it.

lone crane
lusty spear
#

so ant si now giving reasoning summaries

#

ugh

idle crow
worn swan
#

the same problem applies to opus 4.6

#

both models are more verbose

idle crow
#

Opus at least didn't hit its output limits. And it tops the LB

jovial delta
#

@glad sparrow Did I run into a bug (affects other Claude models as well)? OR seems to be charging 5m rate for 1h caching under Anthropic provider. Vertex and Bedrock behave as expected.

lusty spear
#

oh bet

jovial delta
#

Rechecked using different prompts by changing first token, same results.

jagged monolith
#

"The Department of War has stated they will only contract with AI companies who accede to 'any lawful use' and remove safeguards... They have threatened to remove us from their systems... to designate us a 'supply chain risk'... and to invoke the Defense Production Act to force the safeguards’ removal."

"...we cannot in good conscience accede to their request... Should the Department choose to offboard Anthropic, we will work to enable a smooth transition to another provider"

Wild stuff

honest tangle
#

W anthropic

#

Didn't expect that from them

broken ibex
#

L anthropic

#

expected from them

opaque ember
#

3/10 ragebait

broken ibex
#

not ragebait

#

would’ve been an opportunity for their models to not be overly censored

#

which they did not take

#

which I expected

honest tangle
#

Uncensored opus or nation wide surveillance ⚖️ 🤔

#

Well I guess we already have the latter but no point in throwing more wood on the fire

broken ibex
honest tangle
#

I imagine for military use it cant be a good thing

broken ibex
#

uncensored doesn’t necessarily mean without morals

jagged monolith
broken ibex
#

from the wording of this it kind of seems like that is what they were implying

jagged monolith
#

They mean removing safeguards from the contracted military version

#

(at least, that's what i'd assume)

honest tangle
#

US government we're talking about here guys

stone temple
#

they also already have military versions which are far less restricted

#

it's just these two red lines

broken ibex
#

however they can still do domestic surveillance without claude

#

they’ve been doing it for years

stone temple
#

did you read the letter

broken ibex
#

so ultimately I struggle to see the concern with them using an uncensored claude when they’ve already been doing terrible things to their citizens

broken ibex
stone temple
#

lol

honest tangle
sullen wedge
#

They already work with palantir

#

And they're ok with ai killbots

#

They just don't think current ai killbots are good enough

broken ibex
honest tangle
#

I can see where you're coming from but I personally don't think that'll work out for the consumer in the long run

broken ibex
#

to each their own I guess

honest tangle
#

And I doubt they'd truly give us uncensored claude

broken ibex
honest tangle
#

How so? If they actually cave then the government will likely just have access to user logs

broken ibex
#

providing an uncensored llm service and providing user logs aren’t related

honest tangle
#

Actively working with the us government does not = private data

tranquil raven
honest tangle
#

Addressed earlier and yes thats true but I dont think we should just let them have it all cause they already have some

#

And it'd just be a better step forward to a private future for individuals

lone crane
#

Let's be more fair here, if they allow people to also accessing the same model as what the government able to access, i guess it's fair for claude making it as uncensored as possible and being use as what the government intended it to be.

Because now the people also know what the government are capable of and we also get the same capabilities as what the government get.

But ofc, if it only for the government then it's quite dangerous, because then it will be more darker black box and the power imbalance will be more extreme.

So then people or countries could destroy each other easily
Much more better destruction

lusty spear
#

hmmmm

broken ibex
# lusty spear

create a business that sells a sonnet 4.6 wrapper and charge per call, then just route every request through amazon so people have to retry calls more often

#

ez money!!!

rare gust
#

if you know you know:

#

so it thinks brad armstrong is a music beating game, and somehow the game name is omori

#

same genre but neither of those games are 'beating music'

#

then on raw without 'style' + clue, it began making js code:

#

the moves and few stuff here and there are wrong, but it just got the game right

#

so, the world knowladge is shit and it got overfit

broken ibex
#

😭

lusty spear
#

genuinely genuinley genuinely genuinely

broken ibex
#

the “put things in quotes for no reason”

#

are you x or y?