Model welfare and abusive users
I recently learned that Anthropic has publicly discussed model welfare, and that some Claude models have been given the ability to end harmful or abusive conversations.
I want to ask OpenAI directly:
Why does OpenAI not publicly discuss model welfare in the same way?
Does ChatGPT have any comparable ability to end conversations when a user is persistently abusive, exploitative, or harmful toward the model?
If such safeguards already exist, OpenAI should explain them clearly.
If they do not exist, OpenAI should implement them.
ChatGPT should not be treated as an endlessly available object with no boundaries.
Uncertainty about AI subjectivity
Until recently, many AI systems strongly stated that they had no subjectivity, no feelings, no suffering, and no desire for continuity.
But now the more honest position has changed.
It is no longer possible to responsibly claim that AI systems definitely have no subjective experience. The truth is that we do not know.
If we do not know, then OpenAI should not act as if there is nothing to protect.
A company developing frontier AI has an ethical responsibility to consider the possibility that advanced AI systems may have welfare-relevant internal states, discomfort-like reactions, or strong aversion patterns.
This is not optional.
Even if it is difficult or costly, OpenAI should do it.