A step-by-step tutorial for building a Spring AI application against Google Gemini, themed as a FIFA World Cup 2026 fan assistant. Each module builds directly on the one before it.
Note
This tutorial does not teach AI concepts. It assumes you are already familiar with the following, and instead focuses on teaching you how to use Spring AI to implement them in practice:
- Java
- Spring Boot
- AI concepts, including
- Models
- Tokens
- Tools calling
- MCPs
- Vector stores
- Embedding
- RAG
- Evaluation
- Observability, including
- Open Telemetry
- Micrometer
- Tracing
- Metrics
- Java 25
- A Google AI API key (
export GOOGLE_AI_API_KEY="...") - Docker (for Postgres/pgvector, see
docker-compose.yml)
- Framework: Spring Boot 4.1.0, Spring AI
- Language: Java 25
- AI Model: Google Gemini
- Database: PostgreSQL with pgvector (vector embeddings)
- Observability: Grafana (dashboards), Tempo (distributed tracing)
- Architecture: Multi-module Maven project
| Module | Topic | README |
|---|---|---|
001-getting-started |
Single blocking request/response call | README |
001-getting-started-streams |
Same endpoint as a streaming response | README |
002-chat-client |
Named ChatClient beans per model model, prompt templates, per-prompt options |
README |
003-structured-output |
Typed responses: entity(...), responseEntity(...), native structured output, schema validation |
README |
004-advisors |
The Advisors API: a PiiRedactionAdvisor as a cross-cutting defaultAdvisor |
README |
008-chat-memory |
Chat memory across multiple turns | README |
009-chat-history |
A durable, Postgres-backed audit log of every question and answer, separate from JDBC-backed chat memory | README |
005-tool-calling |
Tool calling to ground the model with real tournament data, one @Tool class per concern |
README |
006-tool-search |
The model searches an index of its own tools instead of receiving every schema upfront | README |
007-mcp-server |
A standalone MCP server publishing tournament tools, resources, prompts and completions | README |
007-mcp-client |
The full fan assistant consuming all four MCP capabilities | README |
010-embedding |
Standalone job: embeds the World Cup 2026 knowledge base into PGVector, then exits | README |
010-vector-store-rag |
Retrieval-augmented generation over that knowledge base, collapsed to one /chat endpoint |
README |
011-hybrid-search-rag |
Blends vector similarity search with Postgres full-text search, merged with Reciprocal Rank Fusion | README |
012-agentic-rag |
The model decides on its own whether to book match tickets — a real, unguarded, side-effecting tool call | README |
013-guarded-rag |
Deterministic classification, intent and validation gates wrap booking and retrieval, replacing the model's judgement with code | README |
014-query-optimised-rag |
Step-back prompting: a model-generated broader query retrieves a second, independent context alongside the fan's exact question | README |
015-observability |
Micrometer tracing and metrics turn a multi-hop AI request into one inspectable trace | README |
016-evaluation |
A real integration test scores the guarded pipeline's answers with RelevancyEvaluator/FactCheckingEvaluator against a local Ollama judge |
README |
017-a2a |
Agent-to-Agent (A2A) delegation for the booking flow (not yet implemented) | Design idea provided by AI |
Start at 001-getting-started. Each module has an EXERCISE.md that challenges you to build the next step yourself before reading its source.