TIERRANATURAL DIAMOND
VIEN
Tierra banner 2

DESIGN NOTES

How the agent answers

An English demo of a Vietnamese production support agent. From the client's raw documents to a streamed answer: what was built, and why each choice was made.

6

policy documents

48

searchable chunks

1,307

catalog products

98%

answers in the top 5

01

Architecture

One streaming endpoint, one agent, three kinds of knowledge, each stored the way it is used.

Scroll sideways to see all of itOpen full size
  • POST /api/v1/chat streams text and reasoning in the Vercel AI SDK protocol; tool calls and their raw results stay on the server.
  • The model chooses between two tools itself: there is no hard-coded router.
  • Promotions are not a tool: today's offers are already in the instructions.
  • Every turn is stored in Postgres, and history is rebuilt from those rows, so a client cannot rewrite what the bot said.

02

Preparing the knowledge

A retrieval system repeats whatever its sources say, so the documents were fixed before any code.

  1. 01

    Merged conflicts

    Store phones, resizing and payment rules disagreed across files. Each fact now lives in one place.

  2. 02

    Checked against tierra.vn

    Changed the store list (13 to 17), hours, cash on delivery and buyback terms.

  3. 03

    Instructions out of the knowledge

    Notes written for the chatbot moved to the system prompt instead of being retrieved as facts.

  4. 04

    Translated with trade terms

    Center stone, melee, pavé, GIA grading report. Every number compared automatically.

  5. 05

    Left out what the bot must not say

    The internal staff-training deck, and internal notes in the promotion document.

  6. 06

    One topic per heading

    Two answers were unfindable until a mixed-topic section was split into sub-headings.

DocumentCoversChunks
Company and StoresBrand, collections, 17 stores and hours, hotline8
Products and DesignMetals, colors, ring sizing, customization, production time5
Gemstones and CertificationDiamonds, moissanite, CZ, GIA reports; 25 FAQs16
Ordering and PaymentDeposits, payment methods, card fees, installments5
Delivery and ShippingPickup vs delivery, timelines, inspection, fees6
Warranty and BuybackWarranty, cleaning, buyback rates, resizing, trade-in8

03

Three sources, three storage choices

Not everything belongs in a vector store. Each source is stored the way the agent needs to read it.

Policy documents

Storage
doc_chunks: text + embedding
Read by
Tool call, hybrid search
Why
Free text, asked in many different ways.

Product catalog

Storage
products: one typed row per item
Read by
Tool call, SQL filters
Why
“Under 40 million” and “5.4 mm” are exact values semantic search handles badly.

Promotions

Storage
promotions: offer + date range
Read by
Injected into instructions
Why
They expire, and their conditions must reach the model word for word.

04

Chunking and embedding

The documents are curated and organized by heading, so headings are the natural units. Seven strategies were compared on the same 50 questions.

  1. Step 1

    Split

    at every ### and every numbered FAQ

  2. Step 2

    Merge

    sections under 350 chars with their neighbours

  3. Step 3

    Cap

    split anything over 1,500 chars at paragraphs

  4. Step 4

    Prefix

    the heading path, embedded with the text

  5. Step 5

    Embed

    text-embedding-3-large, 3,072 dims

Ordering and payment > Payment methods

Tierra Diamond accepts cash (VND), bank transfer, card payment (VISA, in store or online) and 0% installment plans on credit cards. Tierra does not accept cash on delivery (COD).
  • The heading path keeps a chunk self-explanatory when it is retrieved alone.
  • A hash of model and text lets unchanged chunks reuse their embedding, so only edits cost an API call.
  • Stored as halfvec(3072): pgvector's HNSW index stops at 2,000 dimensions for plain vector.
  • Result: 48 chunks of about 760 characters.

05

Hybrid retrieval

Two searches with opposite strengths, merged into one ranking.

Scroll sideways to see all of itOpen full size
  • Vector search finds paraphrases: “can I pay when it arrives” for cash on delivery.
  • BM25 finds exact terms embeddings blur: product codes, PT900, 3EX. Accent-stripped, so “bao hanh” matches “bảo hành”.
  • Reciprocal rank fusion merges the rankings without calibrating their scores.
  • No routing by document: a wrong category guess would hide the answer, and 48 chunks gain nothing from it.

06

Monthly promotions

A wrong discount is the costliest mistake this bot can make, so dates are checked in code, never by the model.

Scroll sideways to see all of itOpen full size
  • Marketing's monthly document becomes one markdown file; each offer carries its own valid dates.
  • Filtered in SQL on today's date in Vietnam time: expired offers never reach the model.
  • Two lists, active today and coming up, so “this month?” gets the full answer and “today?” gets the right one.
  • The first live test added 15% + 5% to “20%”; the rules now say stacked discounts compound to 19.25%.

07

The agent

Pydantic AI on FastAPI. The demo runs a free model; the production pick was chosen by measurement on the same question.

ModelVisible reasoningFull answer
nemotron-3-ultra:free (demo)Not measured3.7s
gpt-5.4-mini (production)Summary, from 2.7s3.7s
glm-5.3-flashDefault effort ran out of tokens5.7s
gpt-5.4-nanoNone returned2.6s
gpt-5-miniSummary, from 4.3s12.3s
  • No router: the model reads two tool descriptions and may call both, several times, in one answer.
  • Catalog search is plain SQL. Text matches word by word, and a missing ring type matches any, fixes found by watching real model calls.
  • Guardrails: at most 6 model calls and 4,000 tokens per message, plus per-IP and global rate limits.
  • Each chat gets an access token; the server issues assistant message ids and rebuilds history itself.

08

Tools and how they are used

Three tools, described in their docstrings. Those descriptions are the only routing: the model reads them and decides which to call, with which arguments, and how often.

search_knowledge

Called when
Policies, stores, gemstones, payment, delivery, warranty
Arguments
query
Returns
Top 5 chunks, each with its heading path

search_products

Called when
The customer wants to see or compare products
Arguments
text, category, ring type, stone shape and size, metal, price range
Returns
Products with prices per metal, plus a url and image built in code

escalate_to_staff

Called when
The customer agreed to be connected, or asked for a person
Arguments
reason, summary for staff
Returns
An instruction for the bot's reply, different when staff are off duty
Scroll sideways to see all of itOpen full size
  • Several calls per answer: a budget question with a card-fee question made one knowledge search and two catalog searches.
  • The model picks the arguments: “show me oval engagement rings” became search_products(category, stone_shape) and returned NCH8501.
  • Promotions are not a tool: they are added to the instructions before every model request.
  • Every call is logged on the server and stored with the turn, so a follow-up reuses earlier results instead of searching again.

09

Escalation to staff

Order status, complaints and custom quotes need a person. The bot offers the handoff and escalates only once the customer agrees.

Scroll sideways to see all of itOpen full size
  • Escalation is one-way: from waiting on, the chat route refuses to run the agent.
  • The widget switches to live chat and polls for staff replies; staff reply through a keyed staff API.
  • Off duty, the bot offers a call back and the hotline; on duty, a chat waiting 5 minutes gets one fallback message.
  • Live-chat messages are stored without a model payload, so they never reach the model.

10

Evaluation

50 questions with casual phrasing and typos, each with the exact text its answer must contain, run through the real search.

0.82

right chunk first

0.94

in the top 3

0.98

in the top 5

0.88

mean reciprocal rank

ChoiceCompared withResult
Hybrid searchVector onlyTop 5: 1.00 vs 0.96
text-embedding-3-largebge-m3, text-embedding-3-smallMRR 0.90 vs 0.86 and 0.79
Heading path in each chunkNo prefixTop 5: 0.94 vs 0.88
One topic per headingMixed-topic sectionsTwo questions from unfindable to found

Design comparisons were measured on the Vietnamese knowledge base. Promotions were tested with the live agent on five dates.