An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

2.1.3. FM Customization: Fine-Tuning and Continued Pre-Training

💡 First Principle: Fine-tuning shifts the model's probability distribution toward outputs that match your labeled examples — it changes how the model responds, not what it knows. If your goal is accurate recall of domain facts, fine-tuning is the wrong tool; use RAG. Fine-tuning is the right tool when you need consistent output style, format, or behavior.

When to use each customization approach:
TechniqueWhat It ChangesUse WhenCost Level
Prompt engineeringNothing in the modelBehavior can be guided by instructions$
RAGNothing in the model; adds retrievalNeed factual accuracy about proprietary/recent data$
Fine-tuningModel's output distributionNeed consistent style, format, domain-specific tone$$
Continued pre-trainingModel's knowledge baseNeed the model to deeply internalize a domain corpus — trains on unlabeled domain text, starting from the existing weights (not from scratch)$$

Fine-tuning on Bedrock: Amazon Bedrock supports fine-tuning for select models — currently Amazon Nova, Anthropic Claude 3 Haiku, Meta Llama 3.x, and the Titan image-generation and multimodal-embedding models (Titan Text is not on the list). The process:

  1. Prepare a labeled training dataset in JSONL format (prompt/completion pairs) in S3
  2. Launch a fine-tuning job via Bedrock console or API
  3. Serve the fine-tuned model through Provisioned Throughput or, for the few models that support it (e.g., Amazon Nova), an on-demand custom model deployment (see 2.2.1)
  4. Register the model in SageMaker Model Registry for versioning and lifecycle management

LoRA (Low-Rank Adaptation): For models deployed on SageMaker, LoRA enables parameter-efficient fine-tuning by adding small, trainable rank-decomposition matrices to the frozen model weights. This dramatically reduces compute cost and storage — you train only millions of parameters rather than billions.

Diagnosing a fine-tuned model:
  • Strong on training examples, weak on new queries → overfitting: add more varied examples, train for fewer epochs, and hold out a validation set.
  • Strong on the test set, weak in production → the test set did not represent production traffic (distribution mismatch): evaluate on real, representative queries.
  • General abilities the base model had are now worse → catastrophic forgetting.
  • RLHF needs an input supervised fine-tuning does not: human preference labels ranking pairs of model responses, used to train a reward model.

⚠️ Exam Trap: Fine-tuning does not prevent hallucination. A fine-tuned model still generates statistically plausible outputs and will confabulate facts not in its training data. Combining fine-tuning with RAG is a common production pattern: fine-tune for style/format, add RAG for factual grounding.

Reflection Question: A legal firm wants to build a contract review tool. They have 10,000 labeled examples of contracts with expert annotations marking problematic clauses. They also have 200,000 historical contracts they want the model to "know." What combination of customization techniques would you recommend, and why?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications