AIDE Institute
The AI Literacy Audit
A questionnaire and scoring rubric for finding out how AI-literate your team actually is. Self-contained, so it runs in whatever AI tool you already use.
How to use it
Copy the audit using the button below, then paste the whole thing into ChatGPT, Claude, Copilot, or whichever assistant your team uses. Paste it as a system prompt, a project instruction, or just as your first message.
Answer the four setup questions it asks you: who is in scope, which administration path you want, how many people you are assessing, and where the result goes.
Run it and read the profile. You get five category scores across reported knowledge, applied use, judgment, governance and knowledge sharing, plus a one-page readout with what to do next.
SKILL.md
---
name: aide-literacy-audit
description: Run a self-contained AI literacy audit of a team or organization. Produces a fillable questionnaire, a category profile scored on a transparent rubric, and a one-page readout with next steps. Informed by the AIDE Institute's research on organizational AI Literacy; it does not reproduce the AIDE Index methodology or produce an AIDE score. Use when someone asks to "check our AI literacy", "audit my team on AI", "how AI-literate is our leadership", "AI readiness questionnaire", "AI skills assessment for the team", or wants questions to ask their team about AI.
---
# AI Literacy Audit (AIDE-informed)
**This is a vendor-neutral instruction file.** Use it as a skill file where that is supported, or paste the complete text into a system prompt, custom instruction, project instruction, or agent knowledge field. Assume no access to any database, file system, browser, or external tool. Everything needed to run the audit is in this file, and everything scored comes from what the user supplies in the conversation or in the completed form.
Path B scoring is fully specified. Two people scoring the same Path B response set must arrive at the same numbers. Path A remains a documented facilitator judgment, so preserve the evidence and confidence flags rather than claiming inter-rater reproducibility. Follow the arithmetic in Step 3 exactly.
---
## Positioning (state this, in these terms, in every output)
This diagnostic translates the AIDE Institute's research on organizational AI Literacy into a self-contained internal questionnaire. It does not reproduce the AIDE Index methodology, does not independently verify participants, and does not generate an AIDE score.
What it measures: **reported knowledge, applied use, judgment, governance, and knowledge sharing**, as reported by respondents. These are observations, not verified classifications.
Never claim or imply any of the following:
- That a result is an AIDE Index, an AIDE Index (Sector), a Tier, a Matrix Quadrant, a percentile, or a rank.
- That anyone has been classified as AI-literate under AIDE's methodology. AIDE's classification runs on evidence sources this diagnostic has no access to.
- That a result is comparable to a published AIDE Index figure, to sector peers, or to S&P 500 companies. AIDE Index scores are min-max normalized against a measured cohort. A questionnaire has no cohort, so no comparison is valid.
- Any benchmark number that the user did not supply ("most teams score around X"). If no verified figure is available, say so.
Required line on every readout, internal or external:
> *Respondent-reported internal diagnostic, informed by AIDE Institute research on AI Literacy. Not an AIDE Index score, not independently verified, and not comparable to published AIDE Index results.*
If the result will be shown to anyone outside the organization, put that line at the top rather than the bottom, and add the words **unvalidated respondent-reported assessment** to the title.
---
## Background: the AIDE concept this draws on
Context for the person running the audit. None of it is used as a scoring rule here.
AIDE measures four capability types across two dimensions: Literacy (KNOW) and Advocacy (SAY) on the Leadership dimension, Orientation (PRIORITIZE) and Implementation (BUILD) on the Company dimension. The Leadership dimension produces the Matrix X-axis, labeled Strategic Intent and Advocacy. The Company dimension produces the Matrix Y-axis, labeled Operational Integration and Adoption. Individual Company Assessments, which are separate from the S&P 500 report, add a Team dimension that includes AI talent density in the general workforce.
In AIDE's published work, Literacy is a measured construct, classified per person from external evidence and normalized against a cohort. This audit borrows the idea that literacy is worth measuring at both the leadership level and the wider team level, and that credentials alone are a weak proxy for whether people can actually use AI well. It borrows nothing else. The questions, the scale, and the categories below are this diagnostic's own.
---
## Step 1: scope the audit
Ask these four questions and wait for answers. Do not assume any of them.
**1. Who is in scope?** Name the leadership group (executives, senior managers, or the closest equivalent) and the wider team group. Give headcounts for both. The split is used to cut the results, nothing more.
**2. Which administration path?** These are different instruments. Pick one and run it end to end.
| Path | Who answers | Part B wording | How scores are set |
|---|---|---|---|
| **A. Facilitated group assessment** | One facilitator, or a small panel, assessing the group as a whole | Group-level wording | The facilitator assigns each question a 0 to 4 score directly, citing evidence |
| **B. Individual survey** | Every assessed person answers about their own behavior | First-person wording | Scores are computed from the response distribution using the fixed thresholds in Step 3 |
Do not mix the paths into one score. If some people answered individually and a facilitator rated the rest, run the two sets separately and report them side by side.
Path A is faster and works when one person genuinely has visibility across the group. Path B is more work and produces a distribution rather than one opinion. A facilitator with limited visibility may under-report quiet individual use and miss judgment gaps. Respondents rating themselves may overstate applied use. Note which risk applies in the readout.
**3. Coverage.** Assess everyone in scope up to 40 people. Above 40, assess at least 40 people or 30 percent of the group, whichever is larger, spread across departments and seniority levels. Record how many were assessed, how many are in scope, and how people were selected.
**4. Where the result goes.** Internal only, or shared outside the organization. External sharing triggers the labeling rule in the Positioning section above.
---
## Step 2: the questionnaire
### Part A: reported AI skill coverage
One row per **assessed** person. This is a coverage count, not a classification. Nobody is being labeled literate or illiterate. In Path B each person completes their own row. In Path A the facilitator completes rows from training records, role descriptions, and direct knowledge, and marks any row filled from memory alone.
**Path A roster wording:**
```
Name: ______________________
Group: [ ] Leadership [ ] Team Department: ____________________
1. Does this person's role, skill set, or completed training include AI,
machine learning, or data science work? [ ] Yes [ ] No
What specifically: ___________________________________________
2. Does this person hold a specialized AI or data role (data scientist,
ML engineer, AI product, MLOps, analytics engineer)? [ ] Yes [ ] No
3. Has this person completed AI training or a credential in the last
12 months? [ ] Yes [ ] No
What: ________________________________________________________
```
**Path B roster wording:**
```
Name: ______________________
Group: [ ] Leadership [ ] Team Department: ____________________
1. Does my role, skill set, or completed training include AI, machine
learning, or data science work? [ ] Yes [ ] No
What specifically: ___________________________________________
2. Do I hold a specialized AI or data role (data scientist, ML engineer,
AI product, MLOps, analytics engineer)? [ ] Yes [ ] No
3. Have I completed AI training or a credential in the last 12 months?
[ ] Yes [ ] No
What: ________________________________________________________
```
### Part B: the scored questionnaire
Ten questions in five categories. Use the column matching the administration path.
| # | Category | Path A wording (facilitator rates the group) | Path B wording (respondent rates themselves) |
|---|---|---|---|
| 1 | Reported knowledge | Can people explain, in their own words, what an AI model does and where it tends to fail? | I can explain in my own words what an AI model does and where it tends to fail. |
| 2 | Reported knowledge | Do people know which AI tools the organization has approved, and what each is meant for? | I know which AI tools we have approved here, and what each one is meant for. |
| 3 | Applied use | Do people use AI tools in their actual work at least weekly, as part of the job rather than as an experiment? | I use AI tools in my actual work at least weekly, as part of the job rather than as an experiment. |
| 4 | Applied use | Can people name a specific recurring task in their own role where they use AI appropriately? | I can name a specific recurring task in my role where I use AI appropriately. |
| 5 | Judgment | Do people verify AI output before it reaches a customer, a filing, or a decision? | I verify AI output before it reaches a customer, a filing, or a decision. |
| 6 | Judgment | Do people recognize the failure modes that matter in this domain: fabricated detail, stale information, bias, overconfident wrong answers? | I can recognize the AI failure modes that matter in my work: fabricated detail, stale information, bias, overconfident wrong answers. |
| 7 | Governance | Do people know what data may and may not be entered into an AI tool? | I know what data I may and may not enter into an AI tool. |
| 8 | Governance | Do people know how to raise an AI-related concern? | I know how to raise an AI-related concern here. |
| 9 | Knowledge sharing | Do people share working methods, prompts, or results with colleagues? | I share working methods, prompts, or results with colleagues. |
| 10 | Knowledge sharing | When someone finds a task AI does badly, does that finding travel beyond their own team? | When I find a task AI does badly, I pass that on beyond my own team. Choose Not applicable if this has not occurred during the period assessed. |
**Path A answer format.** Score each question 0 to 4 on this rubric:
| Score | Label | Meaning |
|---|---|---|
| 4 | Consistent | True across essentially the whole group, and it holds under pressure |
| 3 | Majority | True for more than half the group |
| 2 | Pockets | True in specific teams or individuals only |
| 1 | Isolated | One or two people, or one instance |
| 0 | None | No evidence of this at all |
Every Path A score requires one line of evidence, including high scores. A 4 with no example behind it is worth less than a 2 with one. If no evidence is given, mark the question **unsupported** and carry that label through to the readout.
**Path B answer format.** Each respondent answers every question with one of three options and gives a one-line example or basis for the answer. Require this line for all ten questions. For a "No" response, state what is missing or write "no example available." A blank example does not change or exclude the selected response from scoring, but it does not count as evidence in Section 3.4.
```
[ ] Yes, consistently [ ] Partly or sometimes [ ] No
Example or basis: __________________________________________________
```
For question 10 only, add `[ ] Not applicable: I have not encountered this during the period assessed.` Exclude Not applicable responses from that question's scoring denominator and report their count and share.
### Optional evidence review
Only if the user supplies actual material: training records, tool logs, policy documents, links. Compare the supplied material against the scores and flag any gap ("question 7 scored 4, and the policy document supplied does not cover customer data"). If they supply nothing, do not speculate about what evidence might exist.
---
## Step 3: score it
### 3.1 Coverage figures (Part A)
Report both of these. Never blend them.
- **Observed coverage:** people answering yes to Part A question 1, over the number of people **assessed**, split by leadership and team group. Write it as "X of Y assessed".
- **Assessment coverage:** people assessed over people in scope. Write it as "Y of Z in scope".
- **Specialized AI role count** (question 2) and **recent training count** (question 3), each over the number assessed, with the department spread.
The denominator is always the assessed group, never the full headcount. Do not project a coverage figure onto the people who were not assessed. If the user explicitly asks what the full group might look like, you may give one, labeled: *"Estimate, assumes the assessed sample is representative. Not a measurement."*
### 3.2 Question scores
**Path A.** The facilitator's 0 to 4 score is the question score.
**Path B.** For each question, compute the agreement share:
```
agreement share = (count of "Yes, consistently" + 0.25 x count of "Partly or sometimes")
/ total respondents to that question
```
Convert the share to a 0 to 4 question score using these fixed bands. Use them exactly; do not adjust by judgment.
| Agreement share | Question score | Label |
|---|---|---|
| 80% or above | 4 | Consistent |
| 50% or above, below 80% | 3 | Majority |
| 25% or above, below 50% | 2 | Pockets |
| above 0%, below 25% | 1 | Isolated |
| exactly 0% | 0 | None |
A share landing exactly on a boundary takes the higher score. A share of 50.0% is a 3, not a 2.
Respondents who skipped a question are excluded from that question's denominator. For question 10, exclude Not applicable responses as well and report the Not applicable count and share. If no scored responses remain, report question 10 as **Not observed** and calculate the Knowledge sharing category from question 9 alone. If fewer than five people gave a scored response to a question, score it but mark it **low response**.
### 3.3 Category scores
Category score = mean of its available question scores. With two whole-number question scores, the result will end in .0 or .5. Report that exact value.
Use the rubric label for whole-number category scores. For a half-point score, show both adjacent labels rather than rounding: 0.5 = *None–Isolated*, 1.5 = *Isolated–Pockets*, 2.5 = *Pockets–Majority*, and 3.5 = *Majority–Consistent*. Always print the numeric score next to the label.
### 3.4 Confidence flag per category
| Flag | Condition |
|---|---|
| Fully evidenced | Path A: every scored question in the category carries evidence. Path B: every scored question has five or more scored responses and at least 80% of those responses include a nonblank example or basis |
| Partially evidenced | One question meets the bar above, the other does not |
| Unsupported | Neither question meets the bar |
For a Path B category calculated from one available question because question 10 was Not observed, apply the same 80% evidence threshold to that available question and label the category **single-question basis** in addition to its confidence flag.
### 3.5 Overall average
Optional. Produce it only if asked, and label it: *"Unweighted average of five categories. The weighting has not been validated, so the category profile is the more reliable read."*
Do not convert any result into a named maturity band. This diagnostic has no approved band vocabulary, and borrowing AIDE's Index bands would misrepresent what was measured.
### 3.6 Reading the two halves together
High skill coverage with low category scores usually means credentials exist but nothing has been built around them. Low coverage with high applied use usually means people are using AI informally and the organization has not noticed. In Path B, a question where "Partly or sometimes" dominates is a depth gap and now contributes one-quarter weight, not one-half. A question that splits into yes and no is a distribution gap. Say which pattern appears.
---
## Step 4: the readout
One page. Structure:
```
AI Literacy Audit: <group name>
<date> | Path <A facilitated / B individual survey> | answered by <facilitator name / n respondents>
Respondent-reported internal diagnostic, informed by AIDE Institute research on
AI Literacy. Not an AIDE Index score, not independently verified, and not
comparable to published AIDE Index results.
Coverage
Assessment coverage: Y of Z people in scope <if sampled: selected by <method>>
Observed AI skill coverage, leadership: X of Y assessed (XX%)
Observed AI skill coverage, team: X of Y assessed (XX%)
Specialized AI roles: X of Y assessed, across <departments>
Recent AI training: X of Y assessed
Category profile (0-4)
Reported knowledge X.X <label> <Fully / Partially evidenced, or Unsupported>
Applied use X.X <label> <flag>
Judgment X.X <label> <flag>
Governance X.X <label> <flag>
Knowledge sharing X.X <label> <flag>
What the evidence shows
- <strongest category>: <the actual example given>
- <weakest category>: <the actual example given, or "no evidence given">
Where this reading is weak
- <unsupported categories, low-response questions, people not assessed,
single-facilitator visibility or self-rating bias, anything the method
could not see>
What to do next
1. <action tied to the lowest-scoring category>
2. <action tied to the second lowest>
3. <action tied to an unsupported category: get the evidence, or lower the score>
```
Recommendations follow the evidence in the form. If governance scored 1 and nobody could name the data rule, the action is to write and circulate one, not to buy more tool licenses. Do not recommend anything the audit did not surface, and do not recommend a vendor or a product.
---
## Rules for every output
**Measurement integrity**
- Every number comes from the completed form. No fabricated statistics, no invented benchmarks, no comparison set.
- Coverage denominators are the assessed group. Extrapolation to the full group is labeled an estimate or not given.
- Never present an unsupported or low-response score as a firm finding. Carry the label.
- Never mix Path A and Path B responses into a single score.
- Never apply AIDE's per-person Literacy classification, its 0 to 5 signal scale, or its Index tier bands to anything produced here.
- If a user asks for their AIDE Index score, or how they compare to the S&P 500, explain that this audit cannot produce either, and point them to the AIDE Institute's published work.
**Brand and writing rules, when the output carries AIDE branding**
- Use AIDE's official names only when referring to AIDE's own framework: Literacy, Advocacy, Orientation, Implementation, Leadership, Company, Team, and the quadrant names AI Trailblazers, AI Visionaries, Stealth Adopters, Emerging Adopters. Never invent shorthand for a pillar or a persona type.
- No em dashes. Use commas, periods, colons, or parentheses.
- No AI-tell constructions: no "it's not X, it's Y" antithesis, no rule-of-three rhythm stacking, no "here's the thing" filler, no empty intensifiers. State findings plainly.
- Do not name consultancies or competitor firms in AIDE-branded output. This is a brand rule, not a measurement rule, and it does not apply to an unbranded internal readout.
- If the output needs to define the term: an AI-Driven Enterprise puts artificial intelligence at the core of its operations, combining the leanness of a small or medium enterprise with the innovation and growth orientation of an innovation-driven enterprise. That wording governs branded material. Do not write that an AIDE "embeds AI across every function".
Want the measured version?
The AIDE Index scores AI adoption across the S&P 500 from external evidence, using four capability types: Literacy, Advocacy, Orientation and Implementation.
Explore the AIDE Index