REVIEWS / AI MODELS / OWNER INSIGHTS
🦉 WE READ 391 OWNER COMMENTS
DeepSeek R1: what owners actually say
Owners praise DeepSeek R1's open-source reasoning and benchmark-crushing performance, but hit censorship walls, slow local inference, and missing core features like function calling.
What owners complain about
- Censorship on sensitive topics COMMON
Multiple users report the cloud version outright refuses to answer questions about Tiananmen Square, responding with 'I am sorry, I cannot answer that question.' One user noted the model thought through the answer, then deleted its reasoning and said it couldn't help. Interestingly, smaller locally-run models (7b) sometimes bypass this censorship.
- Slow local inference SOME
Running distilled models locally is significantly slower than competitors. One user reported it took 2 minutes to answer 'what is the highest peak in California' while OpenAI o1 took 6 seconds on the same prompt.
- Small distilled models underperform SOME
The 8B distilled model failed basic logic tasks (e.g., 'give me 5 odd numbers that don't contain the letter e') and produced worse results than expected. One user noted token-level blindness makes glyph-based reasoning especially hard for small models.
- No function calling or structured outputs SOME
DeepSeek R1 does not natively support function calling or structured outputs. Users report this is a known limitation with no specific release date, planned for a future major update.
- Questionable actual reasoning ability COMMON
Multiple commenters argue LLMs don't truly reason and are essentially advanced pattern matchers. References to research showing reasoning models fail on scaled-up versions of simple puzzles (e.g., Towers of Hanoi with 25 discs) suggest pattern recognition breaks down on compositional tasks.
What owners love
- Exceptional math and coding benchmarks
Owners highlight 97.3% on MATH-500, 2029 Codeforces rating, and notably solving 2024 Putnam Exam questions that Claude 3.5 Sonnet, GPT-4o, and o1 could not answer satisfactorily.
- Open-source with visible reasoning traces
Users appreciate that reasoning/thinking output is fully visible, contrasting with OpenAI's encrypted and hidden reasoning traces. One commenter noted this transparency forced competitors to partially change course.
- Industry-disrupting pricing
Commenters praise R1 for slashing prices and forcing the entire industry to respond, calling it 'the second most important release of all time right after the original llama.'
- Pure RL approach without supervised fine-tuning
Technically-minded users find the reinforcement learning approach without SFT fascinating and note it performs especially well on closed-system tasks.
- Honesty about knowledge limits
Users were impressed that even smaller distilled models can acknowledge when they don't know details about obscure data structures rather than hallucinating.
Surprising patterns
- Censorship varies by model size and deployment: the cloud version blocks sensitive China-related topics, but locally-run smaller models (7b) sometimes answer the same questions freely, suggesting the censorship is applied at the cloud service layer rather than baked into the base model weights.
- For coding tasks, some developers still prefer Claude 3.5 Sonnet over R1, finding it makes fewer mistakes and questioning whether the 'reasoning/thinking' process actually adds value for practical code generation.
- The model sometimes generates extensive internal reasoning, then discards it entirely and refuses to answer — users observed the thinking trace being deleted before output on censored topics.
WHO SHOULD SKIP IT
Buyers who need function calling, structured outputs, or reliable performance on smaller self-hosted models should wait — and anyone whose work touches topics sensitive to Chinese government policy will hit hard refusals on the cloud version.
Synthesised from 391 real owner comments across 5 platforms. Every point is grounded in the comments — no marketing, no AI guessing. How we do it →