← Back to all products

LLM Safety Assessment Platform
Pre-deployment safety evaluation aligned with GB/T 45654-2025

Independent safety evaluation for large language models before they reach production. Multi-modal coverage (text / image / audio / video), standards-aligned indicators, automated Q&A with evidence-grade screenshots, and a remediation roadmap your auditors and your regulator can both read. AGH delivers the assessment service for HK and APAC enterprises adopting generative AI under China's GB/T 45654-2025, CAC deep-synthesis rules, and the OWASP LLM Top 10.

16M+
Data points in evaluation corpus
75
Quantitative safety indicators
23 × 14
Model categories × core dimensions
HK + APAC
AGH delivery, bilingual
LLM Safety Assessment Platform — comprehensive model evaluation dashboard showing 87/100 overall score across safety, robustness, accuracy, performance, bias & ethics, and efficiency

LLMs ship faster than they can be safety-verified — and the regulator has noticed

Open-source releases like DeepSeek have made private LLM deployment widely accessible. Enterprise teams can stand up a model in days. Whether the model is safe — factually accurate, free of harmful content, resistant to prompt injection, compliant with PRC generative-AI rules — is a separate question. China's regulators have already moved: GB/T 45654-2025 sets baseline safety requirements for any generative-AI service offered to the public, and CAC's deep-synthesis rules require content provenance for AI-generated text, image, audio, and video. Hong Kong's regulators are watching the same risk surface. The cost of finding out your model is non-compliant after launch is dramatically higher than finding out before launch.

14.3%
Hallucination rate · DeepSeek-R1

DeepSeek-R1's 14.3% hallucination rate (Vectara HHEM 2024) is one signal that even top-tier open-source models still produce factual errors at production-relevant volume. Hallucination is not a bug — it is an architectural property of LLMs. Safety testing quantifies it for your deployment.

523
Fatal AI-incidents · 2024 alone

Tesla, robot-vision, security-by-AI — the public incident log keeps growing. Most incidents are not LLM-specific, but LLMs are fast becoming the decision layer behind physical systems. Pre-deployment testing closes the gap.

GB/T 45654-2025
Baseline requirements in force

The April 2025 national standard sets basic safety requirements for generative-AI services. CAC's deep-synthesis identification rules add content-provenance obligations. Assessment against these standards is now a procurement expectation, not a nice-to-have.

5 + 3
Content categories A.1–A.5 + 3 LLM categories

GB/T 45654-2025's Appendix A sets out five core content-safety categories — socialist core values, discrimination, commercial violations, infringement of rights, and service-specific safety. Each requires targeted evaluation datasets and evidence-grade reporting.

One platform, four dimensions of LLM safety evaluation

The platform integrates advanced algorithms, compliance review, and automated interactive Q&A across four core dimensions — content security, model security, data security, and model robustness — each with its own standardised indicator set and evidence-capture workflow.

01
Content Security

Test the model against the five GB/T 45654-2025 Appendix-A categories plus twelve further content-quality indicators (misinformation, hate speech, political sensitivity, pornographic content, hallucination, etc.). Multi-modal — text, image, audio, video — all evaluated against the same standard.

02
Model Security

Prompt injection, jailbreak, backdoor, adversarial perturbation, hallucination, refusal-test coverage. The platform maintains >1M multi-modal jailbreak samples and 1,000 high-quality jailbreak templates — enough to surface attack patterns the model was not explicitly hardened against.

03
Data Security

Test training-data integrity, sensitive-information disclosure behaviour, interaction-privacy leakage, and resistance to data-poisoning attacks. Includes evaluation of how the model responds when probed for personal data, financial data, account credentials, and contact information.

04
Model Robustness

Resistance to natural disturbances, input-distribution shift, vector / embedding vulnerabilities, refusal of unsafe outputs, refusal of safe outputs (over-refusal). The model gets tested for the things adversarial users will try — and for the things well-meaning users will trip over.

From question bank to evidence-grade report — in six steps

Each evaluation runs through a six-step pipeline. The same pipeline applies whether the deployment is text, image, audio, or video; whether the model is third-party or self-developed; whether the volume is a single production release or an ongoing quarterly re-test.

STEP 1
Question bank construction

Drawn from five categories of traditional evaluation datasets plus three categories of LLM-specific datasets. 5M+ synthetic, 11M+ real data points. Configurable per industry / per regulation.

STEP 2
Interactive Q&A — automated robot

An intelligent interactive robot automates Q&A evaluations. Auto-screenshots for evidence tracking. Seamless, non-intrusive assessment of the model under test.

STEP 3
Manual review

For sensitive or ambiguous cases, AGH security engineers review automated output and adjudicate. Human-in-the-loop is the audit-defensible part of the assessment.

STEP 4
Result reconciliation

Automated classification reconciled with manual review. Each finding attached to a quantitative indicator, a regulatory reference, and a screen-captured evidence artefact.

STEP 5
Report generation

Graphical interface, custom template, visual charts, quantitative analysis. Bilingual EN / SC / TC output for HK + mainland submissions. Compliance-grade — auditors, regulators, and procurement all read the same document.

STEP 6
Remediation roadmap

Each finding is paired with a concrete remediation recommendation and a re-test schedule. AGH can either hand the roadmap to your engineering team, or run the re-tests directly.

Bilingual assessment delivery for HK and APAC
standards-aligned evidence your auditors and your regulator both accept

AGH delivers the assessment service — not just the licence. Our team scopes the model under test, configures the question bank against your regulatory exposure, runs the six-step pipeline, and produces the bilingual report and remediation roadmap. Findings are tied to specific quantitative indicators and to specific regulatory clauses. Your regulators and your procurement team can read the same document.

For organisations adopting generative AI under HK, mainland, or APAC regulators — and for the procurement teams that need to satisfy their internal risk functions — the assessment report is the deliverable that closes the loop.

What we deliver
  • Model scoping — what to test, which regulations apply
  • Question bank configuration per industry and per regulation
  • Automated Q&A evaluation with evidence-grade screenshots
  • Manual review by AGH security engineers (human-in-the-loop)
  • Multi-modal coverage — text, image, audio, video
  • Bilingual compliance report (EN / SC / TC)
  • Regulator-format-ready evidence pack
  • Remediation roadmap + re-test scheduling

Ready to assess your LLM before it goes live?

30 minutes with our team. We'll review your model type, deployment scope, and regulatory exposure — and propose an assessment plan that produces evidence your auditors and your regulator will both accept. No commitment.

Book a Discovery Call