New Benchmark Pits Human Writers Against 24 LLMs Across 475 Prompts, Results Show That Only Frontier Models Barely Edge Past Amateur Individuals, Showcasing A Major Skill Gap
Wccftech
A new benchmark by Vulsar AI compared the performance of 24 large language models (LLMs) against human writers using 475 prompts. · The results showed that only the most advanced…
Insights: This benchmark is notable as it underscores the current limitations of AI in creative writing, suggesting that while AI can assist, it is not yet a replacement for skilled human…
AIwritingbenchmark