Part 1 · The Model Under the Product

Chat Template + SFT

Jinja formatting, instruction fine-tuning — LLMs finally learn to talk

THE QUESTION THIS PAGE ANSWERS

ANSWER FIRST

What is the key idea behind “Chat Template + SFT”?

Jinja formatting, instruction fine-tuning — LLMs finally learn to talk

DECISION RULE

Make the claim earn its place. Use this page as a decision aid, not a definition to memorize. Connect the idea to one real task, one observable result, and one failure that would change your mind.

TRY NEXT

Write one question you could answer with evidence after trying this idea.

WATCH FOR

A conclusion that sounds complete but leaves the key assumption untested.

OpenAI agreed on a format, then trained the model specifically on that format, and large models could genuinely hold a conversation
The Evolution from Completion to Conversation
1
Standardize the Conversation Format (Chat Template)
Borrowing from the Jinja template language, define special tokens: <|im_start|> / <|im_end|> to wrap each message
2
Assemble All Messages in the Format
Three roles — system / user / assistant — concatenated into a single flat string and sent to the model
3
SFT (Supervised Fine-Tuning)
Continue training on massive "formatted conversation + high-quality answer" data. The model learns to act as an assistant within this format rather than just completing text.
4
Scale Up
Bigger models, more data: GPT-3 → ChatGPT → GPT-4 — exponential improvement in capability
Chat Template vs. SFT Comparison
Click to switch perspectives
<|im_start|>system← system prompt role marker You are a helpful AI assistant.← System Prompt content <|im_end|>← end marker
<|im_start|>user← user message begins Who is Zixia Fairy? (紫霞仙子是谁?) <|im_end|>
<|im_start|>assistant← model begins completion here (model completion area)← SFT trains the model how to output here <|im_end|>
Before SFT (Base Model)
User: Who is Zixia Fairy? (紫霞仙子是谁?)
紫霞仙子是哥哥,哥哥你喜欢我,你说你喜欢我…
(Continues in the style of training data — nothing like answering a question; Chinese example preserved intentionally)
After SFT (Chat Model)
User: Who is Zixia Fairy? (紫霞仙子是谁?)
Zixia Fairy (紫霞仙子) is a character from the film A Chinese Odyssey, played by Athena Chu. She is Supreme Treasure's destined love, who sacrificed everything for him.
(Understands what answering means; responds as an assistant)
SFT (Supervised Fine-Tuning) is still Token prediction at its core — the only change is that the training data becomes "formatted conversations + high-quality answers".
The model learns to produce a proper response after <|im_start|>assistant, instead of just continuing the training corpus.
At this moment, the "completion machine" evolved into a "conversational assistant"

How “The Evolution from Completion to Conversation” changes an answer

“Jinja formatting, instruction fine-tuning — LLMs finally learn to talk” shows that a model does not process the “word count” we see. It processes Token pieces. Tokenization affects input length, how much context fits, and how much computation a request consumes.

Length, information, and context are different

As “Jinja formatting, instruction fine-tuning — LLMs finally learn to talk” grows, separate three questions: how many Tokens the text becomes, which pieces can change the current decision, and whether older material has fallen outside the context window. Removing repetition is often more useful than simply making the window larger.

Keep what can change the decision

Use “Jinja formatting, instruction fine-tuning — LLMs finally learn to talk” as an A/B test: keep the same question while removing repeated background, compressing format, and trimming irrelevant history. Compare answer quality, latency, and Token count.

From “The Evolution from Completion to Conversation” to “Click to switch perspectives”

“The Evolution from Completion to Conversation” grounds the problem in “1 Standardize the Conversation Format (Chat Template) Borrowing from the Jinja template language, define special tokens: / to wrap each message 2 Assemble All Messages in the Format Thre…”. “Click to switch perspectives” then moves it toward “Chat Template Before vs. After SFT ↺ Replay system ← system prompt role marker You are a helpful AI assistant. ← System Prompt content ← end marker user ← user message begin…”. Together, they show that the lesson is not just a conclusion to remember, but a claim with conditions.

Carry the judgment into the next situation

For long text, keep what can change the conclusion before compressing format and history. A larger context is worth its cost only when the added information is useful.

  • “The Evolution from Completion to Conversation”: 1 Standardize the Conversation Format (Chat Template) Borrowing from the Jinja template language, define special tokens: / to wrap each message 2 Assemble All Messages in the Format Thre…
  • “Click to switch perspectives”: Chat Template Before vs. After SFT ↺ Replay system ← system prompt role marker You are a helpful AI assistant. ← System Prompt content ← end marker user ← user message begin…

The final “Finish by testing the claim” brings the discussion to “Jinja formatting, instruction fine-tuning — LLMs finally learn to talk”. The useful thing to carry forward is knowing which judgments must be revisited when input, scale, or risk changes.

Mark as learned Your reading progress updates automatically
← PreviousNext →

Keep reading

The next useful article in the thread.

ARTICLE DISCUSSION

Leave one useful thought here.

Keep the idea that clicked, the question that stayed open, or a small note for the next learner.

Discussing Chat Template + SFT The Model Under the Product
3discussionsArticle discussion · synced with the Circle
View in the learning circle
AM
Asha MorganContent editor
INSIGHTField note

I turned one judgment from this article into a small experiment I could run today. Knowing what to observe next is more useful than simply remembering the conclusion.

ARTICLE DISCUSSION7 helpful
LH
Lin HarperIndie developer
INSIGHTInsight

After reading this, I first looked for the conditions behind the idea instead of copying the method into a project. That order made the later trade-offs much clearer.

ARTICLE DISCUSSION5 helpful
KM
Kiki MooreProduct operations
QUESTIONQuestion

When this judgment reaches real work, which constraint should be added first? I am curious which step matters most between reading and the first practical attempt.

ARTICLE DISCUSSION4 helpful