Are Your AI Guidelines Any Good?
New research on guidelines for evaluating AI guidelines
AI guidelines are one of those governance techniques that sound intuitively good. Of course we need guidelines, right? They help articulate what responsible people should do when they’re developing AI systems and set expectations for regulators or auditors who want to check whether system behavior is out of bounds and needs to be held accountable. They also help translate more abstract principles into actions that people can and should take. But how do we know if a set of guidelines is actually any good?
A new research paper published at the 2026 Conference on Fairness, Accountability, and Transparency (Pelloth and Zweig, 2026) takes a crack at evaluating the quality of AI guidelines documents. As the scope of their work the authors define an AI guideline as “a document containing recommendations aimed at aiding individuals, organizations, or society in the development of AI systems which have no legal binding.” Their specific focus is on AI development guidelines rather than application or use guidelines, which are important areas for future work.
To develop their evaluation rubric they adopt and adapt material from a medical clinical practice guideline evaluation called AGREE II. The final rubric includes 28 items across 7 domains: scope and purpose, stakeholder involvement, rigor of development, completeness, clarity, applicability, and editorial independence.
To apply the rubric to an AI guidelines document, the authors use a 7-point scale to grade each item. Rather than just average everything together, each domain gets its own normalized score based on the items in that domain. You can see the appendix of the paper for the full ratings manual on how to use it and how they define and rate each of the items. Based on their tests in applying the rubric, they recommend that at least two independent raters be used.
The authors applied the tool to four prominent guidelines — the OECD’s Recommendation of the Council on Artificial Intelligence, UNESCO’s Recommendation on the Ethics of Artificial Intelligence, the EU’s HLEG Ethics Guidelines for Trustworthy AI, and IEEE’s algorithmic bias standard (7003-2024). Each scored poorly on Rigor of Development because they don’t disclose how their recommendations were actually derived.
This is a useful step towards helping to understand the quality of AI guidelines. But what’s lacking is a full-scale evaluation to connect the dots between a low or high-rated guidelines document and its efficacy in practice. Still, it can be a useful tool, not only retrospectively, but also for orienting guideline writers towards aspects that should improve quality.
References
Pelloth JM and Zweig KA (2026) Developing an Assessment Tool for AI Guidelines. Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency: 4511–4551.

