REVIEWS / AI MODELS / OWNER INSIGHTS
🦉 WE READ 122 OWNER COMMENTS
OpenAI o3: what owners actually say
o3's benchmark credibility is under fire, and some owners report it hallucinating worse than cheaper models on practical tasks
What owners complain about
- FrontierMath benchmark scandal COMMON
OpenAI secretly funded Epoch AI's FrontierMath benchmark and had access to problems and solutions before o3's launch, while Epoch was contractually barred from disclosing the relationship. Multiple commenters feel the benchmark results are now meaningless, with one saying 'I genuinely thought the FrontierMath results meant something real.' A verbal agreement that OpenAI wouldn't use the data for training is widely mocked as worthless.
- Hallucinates worse than 4o on real tasks FEW
A side-by-side comparison on a text categorization task showed o3 hallucinated half the semantics and linked concepts in ways it wasn't instructed, producing completely unusable output. The cheaper 4o model got part of the semantics wrong but produced a usable template for ~10 cents, while o3 drained significantly more credits for worse results.
- High cost per prompt SOME
Users report costs around a dollar per prompt for o-series models, making routine use expensive compared to 4o for similar or better practical results on everyday tasks.
- Closed model, cannot self-host SOME
Users explicitly favor DeepSeek and open-weight alternatives because they can be downloaded and run locally, with commenters stating OpenAI's closed approach makes them 'inferior' for anyone who needs local deployment or data sovereignty.
- Requires constant human review for confabulations SOME
Even supporters note you need a review-competent human in the loop to catch confabulations, with best models producing errors within 3-4 paragraphs even in well-documented topics.
What owners love
- Better than 4o for coding
Users report o-series models (o1, o3) are noticeably better than 4o for coding tasks, with o3 appearing improved over 4o at first glance on programming work.
- Strong on complex reasoning benchmarks (credibility disputed)
The model demonstrates impressive results on FrontierMath and other benchmarks, though multiple commenters now question whether those results are trustworthy given the undisclosed OpenAI-Epoch partnership.
- Useful for parsing and categorizing human text
One user reports 4o doing a very good job agreeing with human evaluators on text categorization tasks, and the o-series is seen as a step up for similar structured data work.
Surprising patterns
- A direct head-to-head test found o3 producing worse, unusable output than 4o on a categorization task while costing significantly more — suggesting o3 may not be the default upgrade for every workload.
- The FrontierMath benchmark that o3's headline results depended on was partially funded by OpenAI with data access, and Epoch AI researchers were contractually silenced about it — multiple mathematicians and reviewers unknowingly lent credibility to what some call a 'marketing stunt.'
- Some technically-minded users are abandoning OpenAI entirely for DeepSeek specifically because of open-weight availability, not capability — local hosting matters more to them than benchmark scores.
WHO SHOULD SKIP IT
Anyone who needs verifiable, trustworthy output without human review, or who requires self-hosted models — the cost, confabulation rate, and benchmark credibility concerns make o3 a risky default.
Synthesised from 122 real owner comments across 4 platforms. Every point is grounded in the comments — no marketing, no AI guessing. How we do it →