Field Notes
AI Engineering
14 min· 18 August 2026

RAG in C#: What Actually Separates a Demo from Something You Can Ship

By Ganesh

Every RAG prototype demos beautifully. You load a few dozen documents, ask a few questions you already know the answers to, and it works. Then someone from the actual business asks about invoice INV-2024-8871, and the system confidently describes a completely different invoice — because nothing in your pipeline understood that the number was the whole question.

That gap between the demo and the thing you can put in front of staff is almost entirely about retrieval. Generation is close to a solved problem: give a competent model good context and a clear instruction and it will produce a decent answer. Give it the wrong three paragraphs and it will produce a fluent, confident, wrong one. Most teams spend weeks tuning prompts when the prompt was never the issue.

The short version

  • →Pure vector search fails badly on exact identifiers — hybrid keyword + vector is the default, not an optimisation.
  • →Chunk on document structure, not character counts. A clause split in half retrieves as nonsense.
  • →Reranking does more for answer quality than almost any prompt change.
  • →Build an evaluation set of question → expected-source pairs before you tune anything.
  • →Citations and refusal are trust features. Without them nobody uses it twice.

Chunking, and the mistake everyone makes first

The obvious approach is to split documents every 1,000 characters with some overlap. It's what most tutorials show, it takes four lines, and it quietly destroys your retrieval quality.

The problem is that documents have structure and character counts don't respect it. A policy clause gets cut in half. A table loses its header row, so the numbers in it become meaningless. A heading ends up orphaned at the tail of the previous chunk, so the section it introduces has nothing identifying what it's about.

Split on structure instead — headings, sections, table boundaries — and carry the heading trail into each chunk's text so a fragment retrieved in isolation still knows where it came from. A chunk that begins "Refunds are processed within 14 working days" is far less useful than one that begins "Returns Policy › Refunds › Refunds are processed within 14 working days."

Where the time actually goes

On most engagements, more effort goes into getting documents into a clean, consistently chunked state than into everything else combined. It is unglamorous, it is rarely what the kickoff meeting is about, and it is the difference between a system people use and one they abandon.

Hybrid search, or why your product codes disappear

This is the single highest-value thing in this post, so it gets the most space.

Vector search works by semantic similarity. "How do I get my money back?" and "refund process" land close together in embedding space even though they share no words, which is the entire appeal. But that same property is why INV-2024-8871 fails: an embedding model maps it to roughly the same region as every other invoice number, because to the model they are semantically near-identical. It has no notion that the exact string matters.

Query typeVector onlyKeyword onlyHybrid
Conceptual questionStrongWeakStrong
Exact identifier or SKUPoorStrongStrong
Error code lookupPoorStrongStrong
Paraphrased questionStrongPoorStrong
Person or product nameMixedStrongStrong

There is no column where hybrid is the wrong choice, which is why it should be your starting configuration rather than something you reach for after complaints. Azure AI Search runs both retrieval paths and fuses the results, so this is configuration rather than architecture.

Wiring it up in C#

Two packages do the work: Azure.AI.OpenAI for embeddings and chat, and Azure.Search.Documents for retrieval. Both SDKs have gone through significant surface changes, so pin your versions and check the current reference rather than trusting any snippet you find online, this one included.

csharp
using Azure;
using Azure.AI.OpenAI;
using Azure.Search.Documents;
using Azure.Search.Documents.Models;
using OpenAI.Embeddings;

public sealed class RetrievalService
{
    private readonly EmbeddingClient _embeddings;
    private readonly SearchClient _search;

    public RetrievalService(AzureOpenAIClient openAi, SearchClient search, string embeddingDeployment)
    {
        _embeddings = openAi.GetEmbeddingClient(embeddingDeployment);
        _search = search;
    }

    public async Task<IReadOnlyList<Passage>> RetrieveAsync(
        string question,
        string[] allowedGroupIds,
        CancellationToken ct = default)
    {
        // One embedding call for the question; the documents were embedded at index time.
        var embedding = await _embeddings.GenerateEmbeddingAsync(question, cancellationToken: ct);
        var vector = embedding.Value.ToFloats();

        var options = new SearchOptions
        {
            Size = 20,
            // Semantic reranking runs over the fused hybrid result set.
            QueryType = SearchQueryType.Semantic,
            SemanticSearch = new SemanticSearchOptions
            {
                SemanticConfigurationName = "default-semantic-config"
            },
            VectorSearch = new VectorSearchOptions
            {
                Queries =
                {
                    new VectorizedQuery(vector)
                    {
                        KNearestNeighborsCount = 50,
                        Fields = { "contentVector" }
                    }
                }
            },
            // Security trimming — see below. Never filter in application code.
            Filter = BuildGroupFilter(allowedGroupIds)
        };

        options.Select.Add("id");
        options.Select.Add("title");
        options.Select.Add("content");
        options.Select.Add("sourceUrl");

        // Passing searchText alongside the vector query is what makes this hybrid.
        var response = await _search.SearchAsync<SearchDocument>(question, options, ct);

        var passages = new List<Passage>();
        await foreach (var result in response.Value.GetResultsAsync())
        {
            passages.Add(new Passage(
                Id: result.Document["id"].ToString()!,
                Title: result.Document["title"].ToString()!,
                Content: result.Document["content"].ToString()!,
                SourceUrl: result.Document["sourceUrl"]?.ToString(),
                RerankerScore: result.SemanticSearch?.RerankerScore));
        }

        return passages;
    }

    private static string BuildGroupFilter(string[] groupIds)
    {
        var quoted = groupIds.Select(g => $"'{g.Replace("'", "''")}'");
        return $"groupIds/any(g: search.in(g, {string.Join(",", quoted.Select(q => q))}))";
    }
}

public record Passage(
    string Id,
    string Title,
    string Content,
    string? SourceUrl,
    double? RerankerScore);

The important line is the last search call: passing the question as searchText and supplying a vector query is what makes this hybrid. Drop the first argument and you are back to pure vector search and back to losing invoice numbers.

Retrieve wide, then narrow

Notice the asymmetry: fifty nearest neighbours requested, twenty results returned, and in practice only the top four or five passages go into the prompt. That shape is deliberate. Vector similarity is good at producing a roughly relevant pool and mediocre at ordering it. The semantic reranker is a different and more expensive model that reads the query and each passage properly, and it reorders that pool far better than cosine distance can.

Turning on reranking usually produces a bigger jump in answer quality than any amount of prompt engineering. If you have limited time to improve a RAG system that already works, this is where to spend it.

Making it say "I don't know"

A model handed weak context will still answer. It has no way to tell that the passages are irrelevant, and its training pushes it toward producing something helpful-sounding. You have to make refusal both permitted and easy.

csharp
const string SystemPrompt = """
    Answer only from the numbered sources below.

    Rules:
    - Every factual claim must cite its source as [1], [2], and so on.
    - If the sources do not contain the answer, reply exactly:
      "I couldn't find that in the available documents."
    - Do not use general knowledge. Do not infer beyond what is written.
    - If sources disagree, say so and cite both.
    """;

var context = string.Join("\n\n", passages.Select((p, i) =>
    $"[{i + 1}] {p.Title}\n{p.Content}"));

var chat = openAi.GetChatClient(chatDeployment);

var completion = await chat.CompleteChatAsync(
    [
        new SystemChatMessage(SystemPrompt),
        new UserChatMessage($"Sources:\n\n{context}\n\nQuestion: {question}")
    ],
    new ChatCompletionOptions { Temperature = 0.1f },
    ct);

Two things make this work beyond the instruction itself. Give the model an exact refusal string, because "say you don't know" produces a different apologetic paragraph every time and you can't detect it programmatically. And apply a reranker score floor before you build the prompt at all — if the best passage scores poorly, return the refusal yourself and skip the model call. That's cheaper, faster, and more reliable than hoping the instruction holds.

Measure the refusal rate

A system that never refuses is not confident, it is broken. Track the proportion of questions that end in a refusal and watch it over time. A sudden drop usually means someone loosened a threshold, and it will show up as hallucinations before anyone connects the two.

Security trimming belongs in the query

Filter by permission in the search request, not after results come back. This sounds obvious and is violated constantly, usually by well-meaning code that fetches ten passages and drops the ones the user can't see — which silently shrinks the context, and leaks information through timing and result counts besides.

Store the permitted group identifiers on each document at index time and filter with search.in as in the example above. The practical difficulty is not the filter; it's keeping those group identifiers current when someone changes team. Reindex on permission change, and treat a stale ACL as a defect rather than an inconvenience.

You cannot tune what you cannot measure

This is the step people skip, and skipping it means every subsequent change is guesswork dressed as engineering.

Before touching chunk sizes or prompts, assemble thirty to fifty real questions from the people who will actually use the system, and for each one record which document should answer it. That's your test set. Now every change has a number attached: did the correct document appear in the top five, or didn't it?

Retrieval accuracy is the metric that matters, and it's separable from answer quality. Measure it independently, because when an answer is wrong you need to know whether the right passage failed to surface or the model mishandled a passage it had. Those are different bugs with different fixes, and without the split you will fix the wrong one.

On Semantic Kernel

Every .NET RAG article eventually reaches for an orchestration framework. Semantic Kernel is well built, and it earns its place when you have genuine agentic behaviour — tool selection, planning, multi-step loops where the next call depends on the last result.

Straightforward retrieval is not that. It is one search call and one chat call in sequence, which is about eighty lines of C# you can read end to end and debug with a breakpoint. Adding a framework to that adds indirection between you and the two things most likely to be wrong. My default is to write the pipeline directly and introduce orchestration when the control flow genuinely stops being linear — which, on document question answering, it often never does.

What I'd cut first

A custom chat interface. It is the most visible part and the least consequential, and it competes for time with retrieval quality, which is what actually decides whether anyone comes back after their second question. Put it in Teams, or behind whatever internal tool people already have open, and spend the recovered fortnight on the evaluation set and the reranker.

The other thing worth deferring is conversational memory. Multi-turn follow-ups sound essential and are genuinely hard to do well, because "what about the second one?" has to be rewritten into a standalone query before retrieval can work at all. Ship single-turn, watch how people actually phrase things, then decide whether the complexity is warranted. Frequently they just ask complete questions, because they have learned that complete questions work.

Common questions

Do I need a dedicated vector database?+

Usually not. Azure AI Search handles vectors, keyword search and reranking in one service, and pgvector on a Postgres instance you already run is fine for smaller corpora. Introduce a specialised vector store when you have a measured reason, not by default.

Should I fine-tune instead of doing RAG?+

Almost never for factual question answering over your documents. Fine-tuning teaches style and format, not facts, and it needs redoing whenever the underlying content changes. Retrieval handles changing facts; fine-tuning handles changing voice.

How large should chunks be?+

There's no universal answer, which is why an evaluation set matters more than a number. As a starting point, chunks that hold one coherent idea with some overlap work better than fixed character counts that split sentences.

Is Semantic Kernel required for RAG in .NET?+

No. A straightforward retrieval pipeline is a search call and a chat call, and the SDKs handle both directly. Orchestration frameworks earn their place when you have genuine multi-step agent behaviour, not linear retrieval.

Related engagement

This problem comes up often enough that it has its own entry on the engagements page, with the wider context of how the work is scoped and what typically goes wrong.

See the engagement