The Delegated Authority AI Scorecard

Forty questions in eight weighted categories, with a shortlisting threshold and two elimination gates. We actively suggest you copy it and take it to vendors who are not us.

Read the full method in the buyer's guide
CategoryWeightScore 1 to 5Weighted
A. Evidence and accuracy
%
B. Fit to delegated authority
%
C. Autonomy and controlGate
%
D. Data ownership and privacyGate
%
E. Security and resilience
%
F. Governance and regulatory
%
G. Deployment and integration
%
H. Commercials and exit
%
Total100%
Verdict

Score every category to see a verdict. Gates are applied first: below 3 on C or D is a fail whatever the total says. Otherwise a weighted total at or above 3.5 clears the shortlist bar.

The forty questions

A. Evidence and accuracy

Weight 20%
  1. 1What is your extraction accuracy, on which document types, and against what ground truth?
  2. 2Is accuracy measured continuously in production, or only during implementation?
  3. 3Does every extracted field link back to its source page? Show me.
  4. 4Do you report extraction completeness per submission, so we know how much of a file was captured?
  5. 5What happens when the system is not confident? Does it guess, flag, or stop?
  6. 6Will you publish your accuracy benchmark, or is it only available under non disclosure?

Score 5: accuracy is measured on every correction a user makes, you can see it, and the number moves over time. Score 1: one percentage, no denominator, no date, no method.

B. Fit to delegated authority

Weight 20%
  1. 7How many delegated authority businesses run you in production today, and where?
  2. 8Does the system understand binders, slips, bordereaux, layers, attachment points and cedents natively, or are those just generic documents to it?
  3. 9Can we load our own appetite, wordings and binder terms, and edit them ourselves without raising a request?
  4. 10Can you produce bordereaux in our capacity providers’ formats, including Lloyd’s coverholder reporting standards where they apply?
  5. 11Which languages do you operate in, and is that the whole system or just the interface?

Score 5: named production references, at firms your size, whom you may contact directly. Score 1: “we work with several insurance clients” and no names.

C. Autonomy and control

Weight 10%Gate
  1. 12Can an automated action ever exceed the permissions of the user who triggered it?
  2. 13Which actions need human approval before they take effect, and who decides what goes on that list?
  3. 14How do we set and change the boundaries of autonomous action?
  4. 15What is the rollback path when it does something wrong?
  5. 16Does each agent carry its own identity in the audit record, or does everything appear as the system?

Score 5: permissions inherited from the human, boundaries you configure yourself, sensitive actions gated. Score 1: “the model is very reliable” offered as though it were a control.

D. Data ownership and privacy

Weight 15%Gate
  1. 17Is our data ever used to train or improve any model, yours or anyone else’s? Point me to the clause.
  2. 18Where is our data hosted, and can we choose the region?
  3. 19Is every customer’s data isolated? Could anything we process be seen, inferred or reconstructed by another customer?
  4. 20If we leave, what do we get back, in what format, and how fast?
  5. 21Which sub processors touch our data, and how are we told when that list changes?

Score 5: they point you at a clause. Score 1: “we would never do that”, with nothing in the contract that says so.

E. Security and resilience

Weight 10%
  1. 22Encryption at rest and in transit, backup regime, and recovery objectives?
  2. 23How do users sign in, and does it work with the identity provider we already run?
  3. 24How dependent are you on a single model provider, and what happens if that provider has an outage or changes its terms?
  4. 25Have you been red teamed, by whom, and what did you change afterwards?

Score 5: more than one model provider, infrastructure managed as code, a specific answer on recovery times. Score 1: a security page covered in badges and short on detail.

F. Governance and regulatory

Weight 10%
  1. 26What does the audit trail cover, is it tamper evident, and how long do you keep it?
  2. 27Can you produce, on demand, the full decision trail for one bound risk?
  3. 28Do you provide documentation we can hand to a capacity provider or a regulator, such as a system card?
  4. 29How do you support our obligations under the NAIC Model Bulletin in the states that have adopted it?
  5. 30What is your position on the EU AI Act, and which obligations do you think apply to this system?

Score 5: a precise account of which obligations apply, which do not, and by when. Score 1: urgency about a deadline they cannot cite correctly. Section 8 will tell you which is which.

G. Deployment and integration

Weight 10%
  1. 31How long until a number on our baseline moves, and will you put that date in writing?
  2. 32What do we have to replace? What stays exactly as it is?
  3. 33Which of our systems do you connect to, and are those live in production today or on a roadmap?
  4. 34Who does the configuration work, and what does our team’s time commitment look like week by week?
  5. 35What does change management look like, and who owns adoption when people go quiet in week six?

Score 5: their team does the shaping, your systems stay put, and the time you owe them is quoted in hours. Score 1: a discovery phase, billed separately, before anything runs.

H. Commercials and exit

Weight 5%
  1. 36What is the pricing basis, and what happens to the bill if our submission volume doubles?
  2. 37What is the total first year cost, including implementation, integration and support?
  3. 38What happens if it does not work? Is there a guarantee, and has anyone ever claimed it?
  4. 39What is the notice period, and what are the exit assistance obligations?
  5. 40Which operational metrics will you be measured against, and will you baseline them before we start?

Score 5: they propose the measurement before you think to ask for it. Score 1: a refusal to baseline, on the grounds that value is hard to quantify.