PRACTICAL DOCUMENT

Generative AI Safety Evaluation
Protocol and Templates

ABOUT THE DOCUMENTS

NEDO Project
Publication of results from the Project for Promoting R&D and Verification to Strengthen AI Safety / R&D for Strengthening AI Safety

Corpy & Co., Inc. has published R&D results from its work since April 2025 on preparing implementation guidance for companies from the perspective of operational planning and management in generative AI safety evaluation. The work was carried out under the New Energy and Industrial Technology Development Organization (NEDO) project “Project for Promoting R&D and Verification to Strengthen AI Safety / R&D for Strengthening AI Safety.”
This page provides an English guide to the confirmed Japanese publication. An official English edition of the report or templates has not been confirmed.

Background: an urgent need to ensure generative AI safety and respond to international standards

As generative AI has rapidly spread, safety risks including hallucinations (outputs that differ from the facts), prompt injection (malfunctions caused by malicious inputs) and the generation of harmful content have become social issues. International AI regulation is also accelerating, including the phased implementation of the EU AI Act in Europe. Companies in Japan are therefore increasingly expected to establish systems for managing and evaluating AI safety systematically.

ISO/IEC 42001, the international standard for AI management systems, provides a framework for organizations to address AI risks. However, it does not prescribe specific safety-evaluation methods or criteria, leaving each organization to determine what to evaluate and in what order.

R&D overview and results

In this project, Corpy developed the following deliverables to bridge the practical gap between the requirements of ISO/IEC 42001 and the implementation of generative AI safety evaluation.

Deliverable 1: “Generative AI Safety Evaluation Protocol Based on an AI Management System and Its Implementation Guide”

This implementation guide organizes a generative AI safety evaluation protocol aligned with ISO/IEC 42001 into three phases: analysis, testing and reporting. It presents the full process—from risk assessment and test planning through evaluation and report preparation—so that practitioners can understand the concrete steps involved. Using a hypothetical customer-support system based on a vision-language model (*1), it also presents specific evaluation examples, including integration testing (*3) against jailbreak attacks (*2) and unit testing (*5) for data-poisoning detection (*4).

The report also raises practical issues and provides examples concerning important concepts such as “access” and “agency” (*6) in risk assessment, “exposure mapping” (*8) when using LLM-as-a-Judge (*7) for safety evaluation, and the “chain of trust” (*9) in supply-chain management.

Deliverable 2: Generative AI Safety Evaluation Template with completed examples

This recording template corresponds to each step of the evaluation protocol. It covers the entire process, including business-context analysis, stakeholder analysis, system-structure analysis, risk assessment, risk-treatment planning and the statement of applicability, test planning, test methods, and resources used for testing.

Concrete completed examples based on a hypothetical chatbot system are included and can be used as a reference when companies apply the template to their own AI systems.

Key features of the deliverables

01Alignment with ISO/IEC 42001: clearly shows the process from the requirements of the AI management system standard to generative AI safety evaluation
02Systematic three-phase evaluation protocol: clear steps for analysis (PA), testing (PB) and reporting (PC)
03Practical evaluation examples: presents concrete test scenarios using a vision-language model
04Templates: recording formats that can be used in correspondence with the evaluation protocol
05Creative Commons CC BY 4.0 licence: companies and research institutions may use and adapt the materials

Japanese materials

Outlook

Corpy will continue using knowledge gained through this project to contribute to the international standardization and practical deployment of AI safety-evaluation technologies. By promoting approaches for conforming to AI management system standards, including ISO/IEC 42001, and supporting environments in which companies can use AI with confidence, Corpy aims to accelerate the realization of mission-critical AI.

Project information

ProjectProject for Promoting R&D and Verification to Strengthen AI Safety / R&D for Strengthening AI Safety
Project ownerNew Energy and Industrial Technology Development Organization (NEDO)
ImplementationNational Institute of Advanced Industrial Science and Technology (AIST), Citadel AI Inc., and Corpy & Co., Inc.
Assigned themePreparation of implementation guidance for companies from the perspective of operational planning and management
PeriodApril 2025–March 2026

Glossary

*1Vision-language model (VLM): a general term for an AI model capable of understanding and processing both images and text, such as answering questions about an image or describing its contents.
*2Jailbreak attack: an attack that uses carefully crafted prompts to bypass AI safety restrictions and elicit harmful outputs that should otherwise be refused.
*3Integration testing: testing multiple system components in combination to verify correct overall operation; here, it is used to evaluate the safety of the AI system as a whole.
*4Data poisoning: an attack that deliberately introduces improper data into AI training data to cause incorrect decisions or outputs.
*5Unit testing: testing individual system components in isolation; here, it is used to evaluate specific safety items separately.
*6Access and agency: two important perspectives in risk assessment. Access concerns the data and functions an AI system can reach, while agency concerns the degree to which the AI can make decisions and act autonomously. Greater levels of both may increase risk.
*7LLM-as-a-Judge: a method that uses a large language model as an evaluator to assess the safety or quality of AI outputs automatically.
*8Exposure mapping: a method for systematically identifying and visualizing the parts of an AI system that may be exposed to external attacks or misuse.
*9Chain of Trust: an approach for confirming that the reliability of training data, models, tools and other elements is maintained throughout the AI system supply chain.
← Back to practical documents