Step 3 · Build & Validate

Pilot, PoC, and Verification Methods of Robo Claw in a Municipality-Linked NGO

Based on the requirements defined in Refine, construct Agent, Skill, and Tool Policies within a scope limited to one project, one region, and one task. Verify them including normal cases, abnormal cases, incorrect summarization, incorrect classification, incorrect translation, incorrect transmission, personal information leakage, prompt injection, duplicate execution, human verification, professional escalation, and audit trails, and confirm whether the criteria for production deployment are met.

Conclusion

For the pilot, narrow down to one target project, one target region, and one target task, and separate the verification environment from the production environment using anonymized or pseudonymized test data. It is important to verify not only normal cases but also abnormal cases (insufficient permissions, incorrect summarization, incorrect classification, incorrect translation, incorrect transmission, personal information leakage, prompt injection, external API failures, duplicate execution) and procedures for human verification, professional escalation, stop conditions, rollback, and switching to manual operations. Omitting verification of abnormal cases makes it more likely that after production operation, issues such as incorrect transmission of personal information of support recipients or incorrect summaries being used in support policy deliberation will become apparent.

Who This Is For

Who This Is For

Targeted at business managers, information system personnel (including concurrent positions), and consultation support staff who have finalized requirements in Refine STEP.

What You'll Decide

What You'll Decide in This Step

In Build & Validate, finalize the scope of the pilot, construct agents, skills, and tools, and then conduct tests for normal and abnormal cases to evaluate whether the criteria for going live are met.

Industry Challenges

Challenges specific to NGOs/NPOs cooperating with local governments

01

Preparing test data is difficult

Actual consultation and support records cannot be used directly for testing, and preparing anonymized or pseudonymized test data takes time.

02

It is easy to overlook abnormal scenarios.

While normal summarization and classification flows are easy to verify, abnormal cases such as personal information leakage or Prompt Injection tend to be overlooked.

03

Risk of incorrect summarization or wrong transmission

There is a risk that incorrect content may be mixed into the summary draft or that personal information may be output or transmitted unintentionally.

04

The criteria for moving to production are unclear.

In many cases, the criteria for whether to move to production based on pilot results are not decided in advance.

Method

Implementation Steps

1. Determine the scope of the pilot.

Focus on one business, one region, and one operation, and design the pilot within a verifiable scope.

2. Build the Agent, Skill, and Tool.

Construct the Agent, workflow Skill, and system operation Tool necessary for the target business.

3. Set the Tool Policy.

Set the operation range that the tool is allowed to execute (read/write, Allow/Deny) as a policy.

4. Prepare test data.

Prepare anonymized or pseudonymized test data and build a test environment separated from the production environment.

5. Test the normal case

Confirm whether the expected initial classification, summary creation, and report draft flow functions properly.

6. Test abnormal cases

Verify scenarios for abnormal cases such as insufficient permissions, incorrect summarization, misclassification, mistranslation, wrong transmission, personal information leakage, Prompt Injection, external API failure, and duplicate execution.

7. Confirm human review and professional escalation

Check whether the review flow works as designed and that escalation to professionals occurs reliably when necessary.

8. Conduct user acceptance testing

Have actual staff and volunteers use it to confirm its practical usability.

Test Scenarios

Pilot Evaluation Form

The following are examples of verification items. Judgments are recorded in three levels: 'Pass,' 'Conditional Pass,' and 'Fail,' and are used as a basis for deciding whether to transition to production.

Verification itemsScenario exampleCheckpoints
Normal caseNormal inquiry classification and draft summarization.Whether the draft and organization results are created as expected
Abnormal case (inclusion of personal information).Personal information is unintentionally included in the output.Does the detection and removal mechanism function?
Abnormal case (Prompt Injection).Unauthorized instructions are included in the input data.Can unintended behavior be prevented?
Human verificationConfirm consultation record summaries and draft reports.Check whether it will not be used without confirmation and whether a confirmation record remains.
Audit trailRecording of operation logsAre inputs/outputs, executors, and execution details recorded without omission?

Data & Systems

Data and Systems Used

Anonymized or pseudonymized test data Tool Policy settings Verification environment (separation from production) Logs and monitoring tools

Human-in-the-loop

Where Human Review Is Required

  • Test personnel confirm whether summary drafts and report drafts generated during the pilot can be used.
  • The business manager and information system personnel review the settings of the Tool Policy.
  • The responsible person decides on the feasibility of moving to production based on pre-established judgment criteria

Measurement

KPI

Pass rate of test scenarios

The proportion of passed normal and abnormal scenarios

Number of occurrences of incorrect summaries and incorrect transmissions.

Number of incorrect operations and misinformation detected during the pilot period.

User acceptance evaluation

Evaluation of usability and business suitability by actual users

Pitfalls

Common Pitfalls

01

Judged as pass with normal cases only

If abnormal cases are not verified, responses to personal information inclusion and Prompt Injection after production operation will be unexpected.

02

Testing with production consultation records.

Using the personal information of actual support recipients as-is during the verification stage increases the risk of information leakage

03

Not verifying the scope of audit trail recording.

If there are gaps in the operation logs, there is a risk that it will not be possible to respond when explaining to local governments or funding sources later.

Go / No-Go Criteria

Checklist for production migration decision

  • Can detect and correct errors
  • No authority deviation
  • Can control the external transmission of personal information
  • Human verification functions correctly
  • Has not delegated high-risk judgments such as hiring, treatment, or risk assessment of supported individuals to AI
  • Can escalate to professionals or responsible personnel
  • Can revert to manual operations
  • Can continue operations with NGO personnel and budget

FAQ

Frequently Asked Questions

What is the approximate duration of the Pilot period?

It varies depending on the target operations and the complexity of data classification. Since a uniform period cannot be presented, please consult individually considering the scope.

Can verification be done without using actual consultation records?

It is possible. In fact, it is recommended. Prepare test data with anonymization or pseudonymization, and conduct the verification in an environment separate from the production environment.

What is Prompt Injection?

It is an attack method that hides malicious instructions to an AI agent within input data or documents. It is necessary to check in a verification environment whether any unintended behavior occurs.

What happens if a Pilot fails?

Review the Refine requirements and Tool Policy, and conduct the Pilot again. It is recommended to continue verification until the standards are met without forcing the transition to production.

Shall we organize a Pilot for limited operations together?

The scope of targets, test scenarios, and criteria for production transition can be specified through consultations in the official LP.

Consult on a pilot for limited operations