TOOLING / AGENCY
Why the Judge Grading Our Work Is Blind

Published by Nilesh Patil on LinkedIn on 29 July 2026. Reproduced here unchanged. Read the original on LinkedIn.
Everyone I know uses AI at work now. And everyone has the same complaint.
AI gives you something close. You cannot say exactly what is missing. So you re-prompt and hope.
Hope is not a process.
Jordan Crawford proved the fix with AutoClaygent. He dry-runs research prompts against a scoring rubric until they pass. We generalized it to everything we do at Prospects Pulse Ltd.
Here is the problem it kills.
First attempts are rarely right. A new API config. A new draft. A new approach.
You usually find out after the credits burn or the send goes out.
The loop, in order -
1. It reads the project context first and builds the scoring criteria from it.
2. Dry run on 5 samples. Never 50.
3. A second agent grades the outputs against a weighted rubric. It never sees how we produced them.
4. One change per iteration. Regrade.
5. Pass 8.0, then scale. Fail five rounds, stop honestly.
Every config that passes becomes a recipe in our shared registry.
The next campaign/approach starts from a proven recipe. Not from zero.
Builders grade their own homework kindly. That is why the judge is blind.
Prospects Pulse Ltd. runs every new tool and deliverable through this gate now.