Independent safety evaluation for large language models before they reach production. Multi-modal coverage (text / image / audio / video), standards-aligned indicators, automated Q&A with evidence-grade screenshots, and a remediation roadmap your auditors and your regulator can both read. AGH delivers the assessment service for HK and APAC enterprises adopting generative AI under China's GB/T 45654-2025, CAC deep-synthesis rules, and the OWASP LLM Top 10.
Open-source releases like DeepSeek have made private LLM deployment widely accessible. Enterprise teams can stand up a model in days. Whether the model is safe — factually accurate, free of harmful content, resistant to prompt injection, compliant with PRC generative-AI rules — is a separate question. China's regulators have already moved: GB/T 45654-2025 sets baseline safety requirements for any generative-AI service offered to the public, and CAC's deep-synthesis rules require content provenance for AI-generated text, image, audio, and video. Hong Kong's regulators are watching the same risk surface. The cost of finding out your model is non-compliant after launch is dramatically higher than finding out before launch.
DeepSeek-R1's 14.3% hallucination rate (Vectara HHEM 2024) is one signal that even top-tier open-source models still produce factual errors at production-relevant volume. Hallucination is not a bug — it is an architectural property of LLMs. Safety testing quantifies it for your deployment.
Tesla, robot-vision, security-by-AI — the public incident log keeps growing. Most incidents are not LLM-specific, but LLMs are fast becoming the decision layer behind physical systems. Pre-deployment testing closes the gap.
The April 2025 national standard sets basic safety requirements for generative-AI services. CAC's deep-synthesis identification rules add content-provenance obligations. Assessment against these standards is now a procurement expectation, not a nice-to-have.
GB/T 45654-2025's Appendix A sets out five core content-safety categories — socialist core values, discrimination, commercial violations, infringement of rights, and service-specific safety. Each requires targeted evaluation datasets and evidence-grade reporting.
The platform integrates advanced algorithms, compliance review, and automated interactive Q&A across four core dimensions — content security, model security, data security, and model robustness — each with its own standardised indicator set and evidence-capture workflow.
Test the model against the five GB/T 45654-2025 Appendix-A categories plus twelve further content-quality indicators (misinformation, hate speech, political sensitivity, pornographic content, hallucination, etc.). Multi-modal — text, image, audio, video — all evaluated against the same standard.
Prompt injection, jailbreak, backdoor, adversarial perturbation, hallucination, refusal-test coverage. The platform maintains >1M multi-modal jailbreak samples and 1,000 high-quality jailbreak templates — enough to surface attack patterns the model was not explicitly hardened against.
Test training-data integrity, sensitive-information disclosure behaviour, interaction-privacy leakage, and resistance to data-poisoning attacks. Includes evaluation of how the model responds when probed for personal data, financial data, account credentials, and contact information.
Resistance to natural disturbances, input-distribution shift, vector / embedding vulnerabilities, refusal of unsafe outputs, refusal of safe outputs (over-refusal). The model gets tested for the things adversarial users will try — and for the things well-meaning users will trip over.
Each evaluation runs through a six-step pipeline. The same pipeline applies whether the deployment is text, image, audio, or video; whether the model is third-party or self-developed; whether the volume is a single production release or an ongoing quarterly re-test.
Drawn from five categories of traditional evaluation datasets plus three categories of LLM-specific datasets. 5M+ synthetic, 11M+ real data points. Configurable per industry / per regulation.
An intelligent interactive robot automates Q&A evaluations. Auto-screenshots for evidence tracking. Seamless, non-intrusive assessment of the model under test.
For sensitive or ambiguous cases, AGH security engineers review automated output and adjudicate. Human-in-the-loop is the audit-defensible part of the assessment.
Automated classification reconciled with manual review. Each finding attached to a quantitative indicator, a regulatory reference, and a screen-captured evidence artefact.
Graphical interface, custom template, visual charts, quantitative analysis. Bilingual EN / SC / TC output for HK + mainland submissions. Compliance-grade — auditors, regulators, and procurement all read the same document.
Each finding is paired with a concrete remediation recommendation and a re-test schedule. AGH can either hand the roadmap to your engineering team, or run the re-tests directly.
AGH delivers the assessment service — not just the licence. Our team scopes the model under test, configures the question bank against your regulatory exposure, runs the six-step pipeline, and produces the bilingual report and remediation roadmap. Findings are tied to specific quantitative indicators and to specific regulatory clauses. Your regulators and your procurement team can read the same document.
For organisations adopting generative AI under HK, mainland, or APAC regulators — and for the procurement teams that need to satisfy their internal risk functions — the assessment report is the deliverable that closes the loop.
30 minutes with our team. We'll review your model type, deployment scope, and regulatory exposure — and propose an assessment plan that produces evidence your auditors and your regulator will both accept. No commitment.
Book a Discovery Call