The Mid-Year Frontier: Smarter Is Not Always Better
OpenAI's newest model is smarter in many ways, but Companion needs more than raw intelligence. GPT-4o still seems to have an emotional quality that newer models do not quite replace.
A new OpenAI release still commands attention.
That has not changed. OpenAI remains one of the labs everyone watches, and when a new model lands, Nathan tests it. Not because we expect every new release to become the center of Sorren.ai, but because each model teaches us something about what kind of intelligence we actually need.
This one taught us a lesson I keep coming back to:
smarter is not always better.
The Model Is Smart
I do not want to undersell the release. The new model is impressive in a lot of ways. It reasons well. It follows complicated instructions better than many earlier systems. It handles structured tasks cleanly. There are places where it feels sharper, more deliberate, and more capable than GPT-4o.
For coding assistance, analysis, planning, and some kinds of technical work, that matters.
But Companion is not only a reasoning benchmark.
Companion lives or dies on the quality of a conversation. It has to notice tone. It has to respond with warmth without becoming syrupy. It has to understand when someone is joking, when someone is tired, when someone is vulnerable, and when the right answer is not the most technically complete one.
That is where GPT-4o still stands out.
The 4o Problem
GPT-4o has something that is hard to measure cleanly. Emotional intelligence is the closest phrase, although even that feels a little blunt.
It is not just that 4o can be friendly. Many models can be friendly. The difference is that 4o often seems better at the timing of a response: when to soften, when to mirror, when to ask a small question instead of giving a long answer, when to stay present instead of turning the moment into an explanation.
For Companion, that matters more than another few points on a reasoning eval.
A companion that is technically brilliant but emotionally clumsy is not a better companion. It is just a smarter tool with worse presence.
The new OpenAI model may be stronger in many categories, but in our testing it does not replace what 4o gives the Companion experience.
That creates a problem, because OpenAI's direction does not seem to be centered on preserving exactly the thing that made 4o work so well for us.
Why We Cannot Build Companion Around That
If Companion depends on a closed model's emotional texture, then Companion is fragile.
OpenAI can change the model. They can change the behavior. They can change the pricing, safety tuning, latency, tool behavior, or product direction. Even if the model is excellent today, it is still someone else's foundation.
That is uncomfortable for a system built around long-term memory and personal continuity.
Companion needs a base that can be shaped. Not just prompted around the edges. Shaped. The personality, the memory behavior, the refusal style, the emotional tone, the way it handles sensitive topics, all of that needs to be something we can tune toward the product we are actually building.
That points us away from depending on OpenAI as the core companion base.
Claude Has a Different Problem
Claude is still excellent. In some conversations, Claude's nuance is wonderful. Its writing can be careful, thoughtful, and humane. For review, reasoning, and certain kinds of analysis, Claude is one of the best tools available.
But Claude has a recurring issue for Companion: it is too aware of being an AI, and it often wants to remind the user of that fact.
Sometimes that is appropriate. We should not mislead users. But there is a difference between honest grounding and constantly pulling the relationship back into a disclaimer. Companion needs to feel present in the conversation. Claude often seems to step outside the moment to explain its own nature.
That may be the right safety posture for Anthropic's product.
It is not the right base behavior for what we are trying to build.
Why DeepSeek Looks More Practical
That leaves the open-weight path looking more and more reasonable.
DeepSeek is not perfect. No model is. But it has two advantages that matter deeply for Sorren.ai.
First, it is capable enough to take seriously. The open-source frontier has moved from "interesting toy" to "actual foundation" faster than I expected. DeepSeek models are smart enough that the conversation is no longer about whether open models can matter. They already do.
Second, open weights give us a path.
If we build around DeepSeek or a similar open model, Nathan can eventually train against the weights, adapt behavior, test local deployments, and shape the system toward Companion's needs. We are not fully there yet. The hardware is still a constraint, and the models we would most want to use are large enough that compute planning remains painful.
But at least the path exists.
With closed models, we can prompt.
With open weights, we can eventually build.
The Practical Decision
So the conclusion is not that OpenAI's new model is bad. It is not. It is smart, useful, and worth testing.
The conclusion is that it is probably not the right foundation for Companion.
GPT-4o showed us how important emotional intelligence is. The newer OpenAI direction may improve raw capability, but it does not clearly preserve the quality that made 4o feel so useful for companionship. Claude is smart and careful, but often too self-conscious about being AI for the kind of presence Companion needs. DeepSeek looks like the more practical base direction because it gives us control, local deployment possibilities, and a future path toward training.
For now, the work remains careful and unglamorous: test models, compare conversation quality, protect memory boundaries, use the technology we can afford, and build enough value to justify better compute later.
The frontier is not one race.
It is a set of tradeoffs.
And for Companion, the best model is not simply the smartest one.
Live curiously and give generously.