REVIEWS / AI MODELS / OWNER INSIGHTS

🦉 WE READ 122 OWNER COMMENTS

OpenAI o3: what owners actually say

o3's benchmark credibility is under fire, and some owners report it hallucinating worse than cheaper models on practical tasks

HACKERNEWS · 75 LEMMY · 36 REDDIT · 10 PRODUCTHUNT · 1

What owners complain about

  • FrontierMath benchmark scandal COMMON

    OpenAI secretly funded Epoch AI's FrontierMath benchmark and had access to problems and solutions before o3's launch, while Epoch was contractually barred from disclosing the relationship. Multiple commenters feel the benchmark results are now meaningless, with one saying 'I genuinely thought the FrontierMath results meant something real.' A verbal agreement that OpenAI wouldn't use the data for training is widely mocked as worthless.

  • Hallucinates worse than 4o on real tasks FEW

    A side-by-side comparison on a text categorization task showed o3 hallucinated half the semantics and linked concepts in ways it wasn't instructed, producing completely unusable output. The cheaper 4o model got part of the semantics wrong but produced a usable template for ~10 cents, while o3 drained significantly more credits for worse results.

  • High cost per prompt SOME

    Users report costs around a dollar per prompt for o-series models, making routine use expensive compared to 4o for similar or better practical results on everyday tasks.

  • Closed model, cannot self-host SOME

    Users explicitly favor DeepSeek and open-weight alternatives because they can be downloaded and run locally, with commenters stating OpenAI's closed approach makes them 'inferior' for anyone who needs local deployment or data sovereignty.

  • Requires constant human review for confabulations SOME

    Even supporters note you need a review-competent human in the loop to catch confabulations, with best models producing errors within 3-4 paragraphs even in well-documented topics.

What owners love

  • Better than 4o for coding

    Users report o-series models (o1, o3) are noticeably better than 4o for coding tasks, with o3 appearing improved over 4o at first glance on programming work.

  • Strong on complex reasoning benchmarks (credibility disputed)

    The model demonstrates impressive results on FrontierMath and other benchmarks, though multiple commenters now question whether those results are trustworthy given the undisclosed OpenAI-Epoch partnership.

  • Useful for parsing and categorizing human text

    One user reports 4o doing a very good job agreeing with human evaluators on text categorization tasks, and the o-series is seen as a step up for similar structured data work.

Surprising patterns

  • A direct head-to-head test found o3 producing worse, unusable output than 4o on a categorization task while costing significantly more — suggesting o3 may not be the default upgrade for every workload.
  • The FrontierMath benchmark that o3's headline results depended on was partially funded by OpenAI with data access, and Epoch AI researchers were contractually silenced about it — multiple mathematicians and reviewers unknowingly lent credibility to what some call a 'marketing stunt.'
  • Some technically-minded users are abandoning OpenAI entirely for DeepSeek specifically because of open-weight availability, not capability — local hosting matters more to them than benchmark scores.

WHO SHOULD SKIP IT

Anyone who needs verifiable, trustworthy output without human review, or who requires self-hosted models — the cost, confabulation rate, and benchmark credibility concerns make o3 a risky default.

7.0/10 GYIBB verdict
Full review → Buy on Amazon →

Synthesised from 122 real owner comments across 4 platforms. Every point is grounded in the comments — no marketing, no AI guessing. How we do it →