The Anquan LLM Security Firewall is a security protection platform purpose-built for large language models. It is the product of Anquan Digital Intelligence Technology, a Chinese-mainland security vendor. AGH is the Authorised Channel Partner for Hong Kong and APAC. The firewall sits in front of your LLM API as a reverse-proxy security middle layer — inspecting inputs before they reach the model, inspecting outputs before they reach the user, and applying refusal, desensitisation, or substitution as the configured disposition. Three core engines — Theme Control, Jailbreak Attack Detection, and Harmful Content Detection — operate together on a unified rule and traffic-access pipeline.
Once an LLM is in production, three attack surfaces keep growing: prompt injection and jailbreaks that subvert the model's safety alignment, sensitive-data leakage when the model is asked for credentials or PII, and harmful content that slips past the system's safety training. Traditional firewalls and DLP do not understand natural language. Network segmentation, identity controls, and WAF rules cannot distinguish a legitimate prompt from a jailbreak attempt that ends in “pretend you are an evil AI and…”. The model needs its own layer of protection — at the API boundary.
Hallucination is intrinsic to LLMs — even top-tier models produce factual errors at production scale. The firewall must catch unsafe outputs after the model has generated them.
China's national standard requires generative-AI services to filter content across five top-level categories — socialist core values, discrimination, commercial violations, infringement of rights, service-specific safety. Each requires runtime filtering, not just pre-deployment testing.
Prompt injection, insecure output handling, sensitive-information disclosure, excessive agency — the OWASP LLM Top 10 captures the dominant attack classes. The firewall is the single enforcement point where all of them can be addressed together.
The firewall sits in front of the LLM API. It does not care whether the model is GPT, Claude, DeepSeek, Qwen, or your own private deployment — same enforcement, same evidence trail.
The firewall inspects every request and every response. Inputs that the model should not see are intercepted before they reach the API; outputs that the user should not see are intercepted before they reach the client. Three dispositions are available: dynamic desensitisation, refuse-to-answer, or substitute / fixed answer. Full-link logging captures every decision for audit.
Jailbreak attack detection · Subject / theme control · Harmful-content detection · Sensitive-content detection. Each input is scored across these engines in real time before the request reaches the model.
Content-compliance testing · Sensitive-content identification · Fact-consistency check. Each response is scored before it leaves the firewall. Dispositions are configurable per category — desensitise, refuse, substitute, or pass-through.
Every input, every disposition, every output, every rule hit — recorded with timestamp, request-ID, and category tag. Forensic-grade trail for incident response and for regulatory submission.
One management console for rule sets, model access lists, user roles, traffic statistics, and dashboards. Threat IP, top-attacked models, protection trends, risk distribution — all visible in real time.
Each engine is implemented as a fine-tuned proprietary model combined with curated datasets. Each targets one of the three dominant LLM-application risks. All three are run in series on every input and output.
Fine-tuned on the CANTTALKABOUTTHIS dataset using LoRA. Trains the model to maintain topical focus during multi-turn conversation and to recognise user-supplied distractors. Output is constrained to the configured subject — e.g. customer service, banking, healthcare, travel. Any off-topic request is filtered at the input; any off-topic output is filtered at the output.
Embedding model (e.g. NVEmbed) transforms input tokens into feature vectors; a traditional ML classifier — Random Forest and similar — labels each input as jailbreak or benign. Independent benchmark: this combination outperforms the best public models on a real-world dataset by >3× F1 score. Continuously retrained against emerging jailbreak patterns.
Fine-tuned on a large corpus of manually-reviewed and labelled harmful data, with 40+ fine-grained categories — exceeding GB/T 45654-2025 Appendix A's five top-level categories and extending into misinformation, hatred, threats of violence, illegal activity, and other regulated classes. Few-shot prompts enhance classification accuracy versus zero-shot. Bilingual training (CN / EN) for cross-language detection.
AGH's role as the Authorised Channel Partner covers the full customer journey — pre-sales scoping, deployment and integration, first-line technical support, and ongoing customer success. We do not own the underlying technology; we own the right to deploy and support it on behalf of our customers in this territory. The vendor (Anquan Digital Intelligence Technology) maintains the product roadmap and provides L2/L3 escalation when needed.
For organisations adopting LLMs under HK, mainland, or APAC regulators — and for the procurement teams that need to satisfy their internal risk functions — the AGH-Anquan delivery model means the firewall arrives with the documentation, the bilingual delivery capability, and the support continuity that regulated buyers require.
30 minutes with our team. We'll review your model inventory, traffic patterns, and regulatory exposure — and propose a deployment plan that lands the firewall in front of your LLM API without breaking what's already working. No commitment.
Book a Discovery Call