DeepSeek V4 Flash Self-Evaluation Filtering Accuracy Rises to 88%, Surpassing Claude Fable 5
Source:
x.com
The Stanford team used DeepSeek V4 Flash to sample 5 answers and had the same model self-evaluate and score them, improving accuracy from 79% to 88% on Terminal-Bench 2.1, surpassing Claude Fable 5, with costs only 1/11 of the competitor's. This method validates the potential of test-time scaling; the relevant framework has been open-sourced and is reproducible.