RAG Systems Explained: Architecture, Components & Implementation

A comprehensive guide to RAG systems. Architecture, core components, pipeline design and implementation considerations for production deployments.

What is a RAG system?

A RAG system (retrieval-augmented generation) connects a language model to your organisation's data so it can answer questions grounded in real documents rather than general training knowledge. It's the most practical AI architecture for business knowledge applications.

If you haven't read our intro piece, start with What Is RAG? for the basics. This article goes deeper into the architecture.

Core components

Every RAG system has these building blocks:

Document processor

Handles ingestion. Takes your raw documents (PDFs, Word files, HTML, databases) and converts them into text. This might involve OCR for scanned documents, table extraction, or stripping formatting.

Chunking engine

Splits processed text into smaller pieces (chunks) that the retrieval system can index and search. Chunk size, overlap, and boundary strategy all affect quality. Too small and you lose context. Too large and you dilute relevance.

Embedding model

Converts text chunks into numerical vectors (embeddings) that capture semantic meaning. Similar concepts end up as similar vectors. Popular choices include OpenAI's text-embedding-3-large, Cohere Embed, and open-source models like BGE.

Vector database

Stores embeddings and enables fast similarity search. When a query comes in, the vector DB finds the chunks most semantically similar to the question. Options include Pinecone, Weaviate, pgvector, and OpenSearch.

Retrieval engine

Orchestrates the search. Converts the user query to an embedding, queries the vector DB, and optionally applies re-ranking, filtering, or hybrid search (combining vector + keyword search).

Language model (LLM)

Takes the retrieved chunks plus the user's question and generates a natural-language answer. GPT-4, Claude, and Llama are common choices.

Response formatter

Structures the output: source citations, table formatting, stripping unsafe content, or converting to the format your application needs.

The full pipeline

Here's the end-to-end flow:

  1. Ingest: Documents → processor → chunking → embedding → stored in vector DB (plus metadata)
  2. Query: User question → embedding → vector DB search → top-k relevant chunks retrieved
  3. Augment: Retrieved chunks + system prompt + user question → assembled as context for the LLM
  4. Generate: LLM produces answer grounded in the retrieved context
  5. Post-process: Add source citations, validate output, apply guardrails, return to user

The bottleneck is almost always retrieval, not generation. If the system retrieves the wrong chunks, the LLM can't save it. Focus your optimisation effort on steps 1–3.

Architecture choices

Chunking strategy

Options range from simple (fixed-size with overlap) to sophisticated (semantic chunking based on topic boundaries). For most business documents, 500–1000 token chunks with 100 token overlap is a solid starting point.

Retrieval method

  • Dense retrieval: Vector similarity only. Works well for conceptual questions.
  • Sparse retrieval: Keyword-based (BM25). Better for exact terms, codes, product names.
  • Hybrid: Combines both. Usually the best choice for business applications.

Re-ranking

After initial retrieval, a re-ranker model scores the top results for relevance to the specific question. Adds latency but significantly improves answer quality. Cohere Rerank and cross-encoder models are popular choices.

Metadata filtering

Tagging chunks with metadata (document type, department, date, access level) lets you filter results before or during retrieval. Critical for multi-tenant systems and access control.

Measuring quality

You can't improve what you don't measure. Key metrics for RAG systems:

  • Retrieval precision: Are the retrieved chunks actually relevant?
  • Retrieval recall: Are we finding all the relevant chunks?
  • Answer faithfulness: Does the generated answer accurately reflect the retrieved content?
  • Answer relevance: Does the answer actually address the user's question?
  • Hallucination rate: How often does the model add information not in the sources?

Build an evaluation dataset early. Real questions from real users with expected answers. Run it after every change.

Implementation considerations

  • Data residency: For Australian businesses, deploy on AWS Sydney (ap-southeast-2) to keep data in-country.
  • Access control: Not all users should see all documents. Implement document-level permissions in metadata.
  • Update frequency: How often does your data change? Design your ingestion pipeline accordingly.
  • Cost modelling: Embedding generation (one-time per document) + vector DB hosting + LLM API calls (per query). Model your expected query volume.
  • Start small: Prove value with one knowledge domain before scaling to the whole organisation.

Key takeaways

  • A RAG system has three main phases: ingest, retrieve, and generate.
  • The quality of your retrieval determines the quality of your answers. Garbage in, garbage out.
  • Chunking strategy, embedding model, and retrieval method matter more than the LLM you choose.
  • Start simple, measure relentlessly, and add complexity only when the metrics justify it.
Kasun Wijayamanna
Kasun Wijayamanna Founder & Lead Developer

Postgraduate Researcher (AI & RAG), Curtin University - Western Australia

View profile →

Meet the person

Written by the person who does the work

This article comes from real projects. If it raises a question about your own system, you can ask the founder directly.

HELLO PEOPLE designs, builds and looks after AI, software, app and data solutions for Australian businesses, with senior expertise on every project and a scope agreed before work starts. For software, that means starting with how your business runs, not with the code.

Since 2007, HELLO PEOPLE has delivered more than 100 projects from Perth for small and medium businesses across Australia: custom software and apps, system integrations, data migrations, reporting and dashboards, and AI that works inside the systems a business already runs.

I lead every engagement myself. I trained in accounting before moving into IT, hold accounting and IT professional qualifications and an MBA, and bring more than 20 years of experience across sales, service delivery, inventory and compliance. I am also a PhD candidate in AI at Curtin University, researching retrieval-augmented generation (RAG), so the technology is always judged by what it does for the business.

  • An old-fashioned service

    Small and boutique. The person who scopes your software is the person who builds it, and the same person is there on launch day.

  • Quick responses

    No ticket queue and no account manager in between. You hear back within one business day, usually sooner.

  • A long-term partner

    The first release is the start, not the end. When you need the next system, integration or report, you call the same person, who already knows your business.

Ask the author

Still have a question?

Ask it here and it comes straight to the founder. No sales call, no obligation, and a real answer even if the answer is that you do not need us.

Kasun Wijayamanna, Founder Kasun Wijayamanna
Founder, replies within one business day

Ready to discuss your project?

Tell us what you're working on. We'll come back with a practical recommendation and clear next steps.

Australian owned and operated

Built here. Your data stays here.

  • No offshore development. Everything is written by our own team in Australia. Nothing is subcontracted overseas.
  • Your data stays onshore. Hosted in Australia, on infrastructure you own, under Australian law.
  • Every state, not just ours. Perth, Melbourne, Sydney, Brisbane, Adelaide and everywhere between.