- 🔭 Currently an AI Full-stack Intern @ MINISO — building the group's internal Agent platform and LLM evaluation & data-flywheel pipelines
- 🤖 I build LLM-powered agent systems and full-stack apps — multi-agent orchestration, hybrid RAG pipelines, and end-to-end services with LangGraph / LangChain, FastAPI / Spring Boot, and Vue.js
- 🧩 Active open-source contributor — Apache SeaTunnel, Mastra, Chroma, ContextGem
- 🎓 M.Eng. student in Computer Science at Guangzhou University (2024–2027, GPA 3.96/5.0); research on image generation & editing with diffusion models (Stable Diffusion, CLIP, DreamBooth)
- 💬 Ask me about LLM agents, RAG optimization, LLM evaluation, FastAPI, Spring Boot, or Vue.js
- 🎯 Open to AI application / full-stack development opportunities (graduating June 2027) · 📍 Guangzhou, China
Building the group's Agent platform (centralized access, auth & governance for departmental agents) and the evaluation + data flywheel for MINISO's "XiaoMing" AI customer service.
- Designed centralized auth & routing on Spring Cloud Gateway — dual ordered
GlobalFilters with leaky-bucket rate limiting, covering OAuth authentication, cross-system token exchange, token validation, and unified downstream response wrapping - Replaced full drop-and-rebuild permission sync with an incremental diff-sync strategy — computing the delta between old and new user sets and writing only changes in batches, cutting DB write volume and row-lock hold time
- Built an RBAC permission system on Spring Security, with reflection-based permission-integrity checks at startup that sync permission codes to the database
- Constructed a ~500-question offline eval set for the AI customer service — covering core business, RAG retrieval, sentiment analysis, and safety compliance — with rule-based + LLM-as-Judge evaluators
- Designed the online badcase reflow loop: unified user- and agent-side feedback collection, dual-mode cross-validation, automatic attribution, and routed fixes flowing back into the eval set — closing the data loop
A multi-agent report-review system for China Southern Power Grid: document parsing → knowledge-base construction → multi-agent collaborative review against historical reports and custom criteria.
- Orchestrated multi-node agent workflows in Dify with automatic report-type detection and type-based routing — turning a manually maintained knowledge base into a fully automated pipeline
- Optimized RAG retrieval with hierarchical Markdown chunking, Milvus Hybrid Search, and a Reranker — lifting QA accuracy from 60% to 88%
- Built a multi-stage document-ingestion pipeline engine — queue-chained stages with independent intra-stage concurrency, enabling cross-stage parallel processing of bulk files
- Engineered an OpenAI-compatible LLM connection pool with multi-model dynamic routing, exponential-backoff retry, token-bucket rate limiting, and streaming
A local-first AI office tool for confidentiality-sensitive archives: long-context document extraction plus CLIP-powered image search.
- Deployed the long-context Qwen2.5-1M model locally via Ollama, with Nginx multi-GPU load balancing supporting concurrent intranet users
- Combined PaddleOCR structured recognition with dynamic prompt templates over the ultra-long context window, extracting structured info from whole documents without chunk-boundary loss
- Designed a multi-level caching strategy — layering OCR intermediate states and extraction results — cutting repeated-query latency from minutes to milliseconds
- Implemented text-to-image and image-to-image search for massive asset libraries based on CLIP vision–language alignment
An AI-powered Q&A platform for pet owners: multi-agent orchestration, hybrid RAG, and knowledge-graph long-term memory — built solo, end to end.
- Multi-agent orchestration — LangGraph Supervisor routing + Specialist executors with parallel execution and ReAct-style tool calling; expert mode decomposes complex questions into a dependency-annotated DAG, layered by Kahn topological sort for same-layer parallel / cross-layer serial execution
- Hybrid RAG — vector + BM25 dual recall fused with RRF; parent–child chunk indexing (precise child matching, context-complete parent blocks); fine-grained reranking via
qwen3-rerank - Knowledge-graph memory — Neo4j-based long-term memory: ontology-constrained LLM triple extraction, two-layer entity dedup, vector + full-text + 1-hop-neighbor hybrid retrieval, LPA community clustering, and four-layer provenance (dialogue → chunk → statement → entity)
- apache/seatunnel (9.6k★) — Pulsar connector declarative-validation migration with 14 factory validation tests (PR #11985, in review)
- mastra-ai/mastra (27.5k★) — fixed MCP
tools/callresponses dropping_meta/ui.resourceUri, mirroringtools/listnormalization so third-party MCP hosts can render MCP Apps, with integration tests (PR #22454, in review) - chroma-core/chroma — migrated the Gemini example off deprecated SDK APIs (PR #7637, in review)
- shcherbak-ai/contextgem (2k★) — reported the Gemma3 vision-capability misdetection, confirmed and fixed in v0.12.1 with an explicit vision-capability option (issue #47 → PR #49)
Languages
LLM & Agents
Backend & Data
Tools & Infra