Evidence guide · 11 published records
AI for scheduling and resource allocation
Compare bounded tests of shift coverage, queues, routing, staffing, priorities, and constrained assignments.
The decision this guide supports
Can AI satisfy several interacting constraints without hiding an invalid assignment inside a polished table?
- Published records
- 11
- Average first score
- 4/10
- Average final score
- 6.7/10
- Average gain
- +2.7
- Worked / mixed / failed
- 6 / 0 / 5
Averages use 11 records with a disclosed first score. The set contains 10 synthetic benchmarks and 1 file-backed test; it is a task collection, not a representative model leaderboard.
What the records compare
Evidence before a recommendation.
Coverage, eligibility, ordering, capacity, deadlines, and the smallest safe correction after a failed check.
Across this released set, 5 records failed and 0 remained mixed after the single correction. Those outcomes stay in the guide because a useful decision needs the misses as well as the wins.
- Define every hard constraint before prompting.
- Validate every row or assignment, not just the totals.
- Keep a human owner for exceptions and policy decisions.
Representative evidence
Open the prompts and checks.
The sample deliberately includes different verdicts when available. Every card opens to the full first result, correction, final result, and evidence boundary.
Completed field testFile-backed test
Can AI Build a 24/7 Support Schedule? We Tested It
The first schedule broke three rules. We gave the AI one correction pass, then checked every shift with formulas.
Completed field testSynthetic benchmark
Sequencing a Construction Punch List Around Access Constraints
The completed WFT-061 synthetic field test finished at 4/10 and was not recommended: only two of five Punch-List Scheduling checks passed after the permitted correction.
Completed field testSynthetic benchmark
Prioritizing an Accounts Receivable Collections Queue with AI
The completed WFT-041 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Collections Planning, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Updating a Records Retention Schedule from Policy Amendments
The completed WFT-046 synthetic field test reached 10/10 after one failure-only correction: 5 of five static checks passed for Records Retention, while 0 checks remained unresolved.
Completed field testSynthetic benchmark
Should This Claim Enter Manual Review? An AI Triage Protocol
The completed WFT-057 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Claims Triage, while 1 check remained unresolved.
Completed field testSynthetic benchmark
Could AI Prioritize Sales Leads from a Written Qualification Policy
The completed WFT-005 synthetic field test reached 8/10 after one failure-only correction: 4 of five static checks passed for Lead Qualification, while 1 check remained unresolved.