{"id":"2089421448784023553","url":"https://x.com/jackyk02/status/2089421448784023553","text":"Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper 💰\n\nAs open-source models become more capable, they can now generate large numbers of high-quality candidate solutions and verify their own outputs at very low cost.\n\nFor example, we find that sampling just 5 solutions with DeepSeek V4 Flash and ranking them using the same model with LLM-as-a-Verifier can lead to a significant boost in accuracy (79% → 88%), outperforming closed frontier models on Terminal-Bench.\n\nTry it out today: https://github.com/llm-as-a-verifier/llm-as-a-verifier#self-verification-terminal-bench-21\n\nMore on verification scaling in my previous post.","author":{"name":"Jacky Kwok","username":"jackyk02","avatarUrl":"https://pbs.twimg.com/profile_images/1938668993293750272/iNSxmD3Q_200x200.jpg"},"createdAt":"Mon Aug 17 18:37:47 +0000 2026","engagement":{"replies":153,"retweets":400,"likes":3118,"views":1005861},"media":{"photos":[{"url":"https://pbs.twimg.com/media/HP8c7DyaYAAjGdm.jpg?name=orig","width":2286,"height":1740}],"videos":[]},"quoteTweet":{"id":"2074969820739805275","url":"https://x.com/jackyk02/status/2074969820739805275","text":"How can we extract richer signals from AI Feedback?\n\nIntroducing LLM-as-a-Verifier✨— a simple verification scaling framework that achieves SOTA on agentic benchmarks 🚀\n\nThe key idea:\n- Use fine-grained scoring granularity (e.g., 1-20 instead of the standard 1-5 scale)\n- Take the expectation over the full logprob distribution of score tokens\n- Scale repeated evaluation and criteria decomposition\n\nYou can use these fine-grained signals for more effective test-time scaling, RL, and agent monitoring! It achieves SOTA across Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench 👑\n\nAdvised by @Azaliamirh @istoica05 @drmapavone @chelseabfinn \n\n🧵👇","author":{"name":"Jacky Kwok","username":"jackyk02","avatarUrl":"https://pbs.twimg.com/profile_images/1938668993293750272/iNSxmD3Q_200x200.jpg"},"createdAt":"Wed Jul 08 21:32:10 +0000 2026","media":{"photos":[{"url":"https://pbs.twimg.com/media/HMvFN6qbMAA6FeR.jpg?name=orig","width":2698,"height":1667}],"videos":[]}},"externalLink":{"url":"https://github.com/llm-as-a-verifier/llm-as-a-verifier#self-verification-terminal-bench-21","displayUrl":"github.com","title":"GitHub - llm-as-a-verifier/llm-as-a-verifier: LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and medical agentic benchmarks.","description":"LLM-as-a-Verifier is a general-purpose framework that provides fine-grained feedback for any agent without requiring additional training. It achieves SOTA performance across coding, robotics, and m...","thumbnailUrl":"https://opengraph.githubassets.com/807815ceed01e72ad9bcf63fd4c11e8105ad1e193f7ee620b58e45e6aba4260d/llm-as-a-verifier/llm-as-a-verifier"},"adhxContext":{"savedByCount":1,"publicTags":[],"previewUrl":"https://adhx.com/jackyk02/status/2089421448784023553"}}