7.5 / 10
GPT-5 mini
307 USER VOICES · REALITY SCORE
Smaller OpenAI model excels at structured tool-calling and agentic tasks but is highly prompt-sensitive, raising questions about benchmark v