The Delegated Authority AI Scorecard
Forty questions in eight weighted categories, with a shortlisting threshold and two elimination gates. We actively suggest you copy it and take it to vendors who are not us.
Read the full method in the buyer's guide| Category | Weight | Score 1 to 5 | Weighted |
|---|---|---|---|
| A. Evidence and accuracy | % | — | |
| B. Fit to delegated authority | % | — | |
| C. Autonomy and controlGate | % | — | |
| D. Data ownership and privacyGate | % | — | |
| E. Security and resilience | % | — | |
| F. Governance and regulatory | % | — | |
| G. Deployment and integration | % | — | |
| H. Commercials and exit | % | — | |
| Total | 100% | — |
Score every category to see a verdict. Gates are applied first: below 3 on C or D is a fail whatever the total says. Otherwise a weighted total at or above 3.5 clears the shortlist bar.
A. Evidence and accuracy
- 1What is your extraction accuracy, on which document types, and against what ground truth?
- 2Is accuracy measured continuously in production, or only during implementation?
- 3Does every extracted field link back to its source page? Show me.
- 4Do you report extraction completeness per submission, so we know how much of a file was captured?
- 5What happens when the system is not confident? Does it guess, flag, or stop?
- 6Will you publish your accuracy benchmark, or is it only available under non disclosure?
Score 5: accuracy is measured on every correction a user makes, you can see it, and the number moves over time. Score 1: one percentage, no denominator, no date, no method.
B. Fit to delegated authority
- 7How many delegated authority businesses run you in production today, and where?
- 8Does the system understand binders, slips, bordereaux, layers, attachment points and cedents natively, or are those just generic documents to it?
- 9Can we load our own appetite, wordings and binder terms, and edit them ourselves without raising a request?
- 10Can you produce bordereaux in our capacity providers’ formats, including Lloyd’s coverholder reporting standards where they apply?
- 11Which languages do you operate in, and is that the whole system or just the interface?
Score 5: named production references, at firms your size, whom you may contact directly. Score 1: “we work with several insurance clients” and no names.
C. Autonomy and control
- 12Can an automated action ever exceed the permissions of the user who triggered it?
- 13Which actions need human approval before they take effect, and who decides what goes on that list?
- 14How do we set and change the boundaries of autonomous action?
- 15What is the rollback path when it does something wrong?
- 16Does each agent carry its own identity in the audit record, or does everything appear as the system?
Score 5: permissions inherited from the human, boundaries you configure yourself, sensitive actions gated. Score 1: “the model is very reliable” offered as though it were a control.
D. Data ownership and privacy
- 17Is our data ever used to train or improve any model, yours or anyone else’s? Point me to the clause.
- 18Where is our data hosted, and can we choose the region?
- 19Is every customer’s data isolated? Could anything we process be seen, inferred or reconstructed by another customer?
- 20If we leave, what do we get back, in what format, and how fast?
- 21Which sub processors touch our data, and how are we told when that list changes?
Score 5: they point you at a clause. Score 1: “we would never do that”, with nothing in the contract that says so.
E. Security and resilience
- 22Encryption at rest and in transit, backup regime, and recovery objectives?
- 23How do users sign in, and does it work with the identity provider we already run?
- 24How dependent are you on a single model provider, and what happens if that provider has an outage or changes its terms?
- 25Have you been red teamed, by whom, and what did you change afterwards?
Score 5: more than one model provider, infrastructure managed as code, a specific answer on recovery times. Score 1: a security page covered in badges and short on detail.
F. Governance and regulatory
- 26What does the audit trail cover, is it tamper evident, and how long do you keep it?
- 27Can you produce, on demand, the full decision trail for one bound risk?
- 28Do you provide documentation we can hand to a capacity provider or a regulator, such as a system card?
- 29How do you support our obligations under the NAIC Model Bulletin in the states that have adopted it?
- 30What is your position on the EU AI Act, and which obligations do you think apply to this system?
Score 5: a precise account of which obligations apply, which do not, and by when. Score 1: urgency about a deadline they cannot cite correctly. Section 8 will tell you which is which.
G. Deployment and integration
- 31How long until a number on our baseline moves, and will you put that date in writing?
- 32What do we have to replace? What stays exactly as it is?
- 33Which of our systems do you connect to, and are those live in production today or on a roadmap?
- 34Who does the configuration work, and what does our team’s time commitment look like week by week?
- 35What does change management look like, and who owns adoption when people go quiet in week six?
Score 5: their team does the shaping, your systems stay put, and the time you owe them is quoted in hours. Score 1: a discovery phase, billed separately, before anything runs.
H. Commercials and exit
- 36What is the pricing basis, and what happens to the bill if our submission volume doubles?
- 37What is the total first year cost, including implementation, integration and support?
- 38What happens if it does not work? Is there a guarantee, and has anyone ever claimed it?
- 39What is the notice period, and what are the exit assistance obligations?
- 40Which operational metrics will you be measured against, and will you baseline them before we start?
Score 5: they propose the measurement before you think to ask for it. Score 1: a refusal to baseline, on the grounds that value is hard to quantify.
