Salesforce reports human oversight where required, but does not publish the score definition, calibration, eligible-case share, or false-action rate.
IntelligenceMake the confidence score earn the authority.
Replay one agent decision against real past cases. Compare autonomy cutoffs, false actions, review load, and cost before the score becomes permission.
No upload is required. Your inputs stay in this browser unless you export them or request a review.
Internal processing only. Bank time did not change.
Customer-supplied result in a vendor release.
A threshold moves error and review work. It does not remove either.
Price the decision, not the software.
The same score should not control a low-risk draft and a high-cost account change. Write the business action, the cost of a wrong move, and the maximum error rate leadership will accept.
Enter what happened after the score.
Use a representative, time-bounded sample from the same task, channel, user group, model, tools, and policy you plan to release. “Correct” means the full business outcome passed, not that the text looked good.
| Confidence band | Cases tested | Fully correct | Wrong actions | Observed accuracy | Band gap |
|---|
See where each cutoff moves the work.
Higher thresholds usually reduce automated volume and wrong actions while increasing human review. The preferred gate must satisfy the accuracy, error, and sample floors before cost is considered.
The delegation decision.
This recommendation is based on the entered historical sample and business limits. It is a release-planning aid, not a platform guarantee, legal opinion, or substitute for controlled production monitoring.
Sources, method, and limitations
- Salesforce and Air India, September 15, 2026. Salesforce says eligible email responses are generated only above a 95% confidence threshold, with human oversight where required. It also reports refund processing fell from about 14 days to four hours and name changes from three days to 30 minutes. These are vendor-published, customer-supplied results. The release does not define the score, calibration method, eligible-case share, false-action rate, or exact human-review rule.
- Google Machine Learning Crash Course, updated January 12, 2026. Different decision thresholds change the balance of true positives, false positives, true negatives, and false negatives.
- Google guidance on ROC and AUC, updated May 11, 2026. The preferable threshold depends in part on the relative business cost of false positives and false negatives.
- scikit-learn probability calibration guidance. A well-calibrated probability near 0.8 should be correct about 80% of the time for comparable cases. Not every classifier or confidence score is calibrated.
- NIST AI RMF Playbook, Measure. NIST recommends measuring override decisions, documenting results, and feeding those findings into continual improvement.
Method: For each candidate gate, the lab uses the historical cases at or above that confidence band. It calculates empirical accuracy, qualifying sample size, scaled monthly auto volume, scaled wrong actions, estimated review cost, estimated wrong-action cost, and their sum. A gate must meet the user-entered accuracy floor, wrong-action ceiling, and sample floor. Unknown score meaning, an irreversible action, a sensitive decision, weak documentation, or a large calibration gap creates an additional release warning. The lowest modeled-cost passing gate is recommended. These are Aule planning rules, not vendor, regulatory, or statistical-certification requirements.