Hello! Welcome to our final lesson on advanced prompting strategies.
In our last lesson, we focused on prompt engineering as a discipline for optimizing a single reasoning path. We learned to craft detailed "specifications" using frameworks like CRISP-E to guide a model toward a high-quality output. This is perfect for tasks where the best path is known.
But what about problems that are inherently exploratory, with many potential dead ends and no single, obvious route to a solution? For these, we need to move beyond a linear chain of thought. Today, we'll explore a powerful strategy that does just that.
Our learning outcome is to apply advanced prompting strategies like Tree of Thoughts for complex problem-solving. Instead of just telling the model what to do, we'll give it a framework to explore, evaluate, and navigate a complex problem space on its own.
This lesson will cover:
- The conceptual leap from Chain-of-Thought (CoT) to Tree of Thoughts (ToT).
- The core components of the ToT framework: thought decomposition, generation, evaluation, and search.
- How ToT is applied to both logical and creative tasks using different strategies.
- A practical look at the prompts and code structure needed to implement ToT.
1. From a Linear Chain to a Branching Tree
In previous lessons, we saw how Chain-of-Thought (CoT) prompting helps models solve problems by generating intermediate reasoning steps. However, CoT is fundamentally linear—it generates one step after another. If the model makes a mistake early on, it's stuck on that path. Self-Consistency improves on this by running multiple independent chains and voting, but it doesn't allow for exploration within a single reasoning process.
Tree of Thoughts (ToT) introduces a new paradigm. It allows a large language model (LLM) to perform a more deliberate, "System 2" style of thinking by exploring multiple reasoning paths simultaneously.

As you can see, ToT structures the problem-solving process as a search through a tree. Each "thought" is a node, and the model can generate multiple potential next thoughts (branches), evaluate them, and decide which paths to pursue, prune, or even backtrack from. This makes it far more robust for complex problems where exploration is key.
2. The Four Pillars of the ToT Framework
ToT isn't a single prompt; it's a framework that orchestrates multiple LLM calls within a classic search algorithm. Your background in computer science will make this structure feel very familiar. Implementing ToT for any given problem requires answering four key questions.
To get a clear overview of how ToT works and its core components, please watch the following segment from Yannic Kilcher's review of the original paper.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models (Full Paper Review)
This video provides a deep dive into the Tree of Thoughts paper. This segment explains the core mechanism, contrasting it with Chain of Thought and detailing how the search process is guided by the LLM itself.
Watch from 06:18 to 10:49. Focus on understanding how ToT uses the LLM for both generating thoughts and for self-critique, enabling it to prune branches and backtrack. Notice the parallel to classic tree search algorithms.
As the video explained, setting up a ToT system boils down to defining these four components:
-
Thought Decomposition: How do you break the problem into intermediate steps? A "thought" needs to be small enough for an LLM to generate diverse options, but large enough to be meaningfully evaluated. For a math problem, it might be one equation. For a writing task, it could be a paragraph outline.
-
Thought Generator: How do you create potential next steps from a given state? There are two main strategies:
- Sample: Generate multiple independent and identically distributed (i.i.d.) thoughts. This is good for creative tasks where diversity is important. You'd make
kseparate API calls. - Propose: Ask the LLM to generate a list of
kdifferent thoughts in a single prompt. This is more efficient for constrained tasks (like math) and helps avoid duplicates.
- Sample: Generate multiple independent and identically distributed (i.i.d.) thoughts. This is good for creative tasks where diversity is important. You'd make
-
State Evaluator: How do you heuristically score the promise of each thought? This is where the LLM acts as a critic.
- Value: Ask the LLM to assign a scalar score (e.g., 1-10) or a class (e.g.,
sure/likely/impossible) to each thought independently. This is useful when progress can be measured objectively. - Vote: Present several thoughts to the LLM and ask it to choose the most promising one. This works well for subjective tasks like writing, where direct scoring is difficult.
- Value: Ask the LLM to assign a scalar score (e.g., 1-10) or a class (e.g.,
-
Search Algorithm: Which algorithm will you use to navigate the tree? The original paper explores two simple, effective options:
- Breadth-First Search (BFS): Explores the tree level by level, keeping the
bbest thoughts at each step. It's exhaustive but can be resource-intensive. - Depth-First Search (DFS): Explores the most promising path all the way to a solution. If it hits a dead end or a low-value state, it backtracks and tries another branch.
- Breadth-First Search (BFS): Explores the tree level by level, keeping the
The real innovation of ToT is using the LLM itself for steps 2 and 3, effectively making the search heuristics language-based and dynamically reasoned.
3. Applying ToT: From Logic Puzzles to Creative Writing
To make this concrete, we'll look at a superb blog post from Hugging Face that implements ToT in Python for two different tasks, showcasing the framework's flexibility. It's an excellent resource that connects the theory directly to code.
Understanding and Implementing the Tree of Thoughts ...
This article, 'Understanding and Implementing the Tree of Thoughts Paradigm', provides a full-fledged Python implementation of the ToT framework. It is an invaluable resource for understanding the practical details.
Read the article from the beginning up to the start of the 'A Reusable TreeOfThoughts Class' section. Pay close attention to: The 'Creative Writing' section: Note how it uses the sample generator and vote evaluator with a BFS search. The 'Game of 24' section: Note how it uses the propose generator and value evaluator with both BFS and DFS search algorithms. Focus on understanding the structure of the prompts and the logic of the Python code, especially how the different components (generator, evaluator, search) come together.
Let's distill the key insights from that implementation.
Case Study 1: Game of 24 (Logical, Constrained Task)
This task requires finding a mathematical expression to reach 24 using four given numbers.
- Thought Decomposition: A thought is a single arithmetic operation (e.g.,
4 + 8 = 12). - Thought Generator: Uses the
proposestrategy. A single prompt asks the LLM to list several possible next operations given the remaining numbers. This is efficient and avoids duplicates. - State Evaluator: Uses the
valuestrategy. A separate prompt asks the LLM to classify the likelihood of reaching 24 from the remaining numbers assure,likely, orimpossible. These are then mapped to numerical scores. - Search Algorithm: Both BFS and DFS are effective. The blog post demonstrates how a BFS maintains the top
b=5paths at each step, while a DFS explores one promising path deeply, backtracking if the value drops below a threshold.
The results are striking: while standard CoT solves only 4% of "Game of 24" tasks, ToT achieves a 74% success rate.
Case Study 2: Creative Writing (Open-Ended, Exploratory Task)
Here, the task is to write a coherent passage where four paragraphs end with four specific, random sentences.
- Thought Decomposition: The process is broken into two main thoughts: (1) creating a high-level plan, and (2) writing the final passage based on that plan.
- Thought Generator: Uses the
samplestrategy. The LLM is promptedk=5times to generate 5 different, diverse plans. This is ideal for creative exploration. - State Evaluator: Uses the
votestrategy. The 5 generated plans are presented to the LLM in a new prompt, which asks it to analyze them and vote for the best one. - Search Algorithm: A simple BFS with a breadth of
b=1is used. After generating and voting on 5 plans, only the winning plan is kept to generate the final passages, which are then voted on again.
This shows how ToT can be adapted to problems where the "correct" answer is subjective, significantly improving passage coherence compared to linear methods.
Test your understanding!
Imagine you're using an LLM to debug a complex bug in a Python application. The bug could be in the frontend (React), the backend (Django), or the database query (SQL). A linear CoT approach might get stuck investigating the wrong component.
How would you design a ToT framework to tackle this? Briefly outline your choices for the four key components.
- Thought Decomposition:
- Thought Generation Strategy:
- State Evaluation Strategy:
- Search Algorithm:
Show answer
Here’s one possible way to structure a ToT approach for debugging:
-
Thought Decomposition: A "thought" could be a specific hypothesis about the bug's location and cause. For example: "Hypothesis: The bug is a race condition in the React frontend's state management." or "Hypothesis: The bug is an incorrect JOIN in the SQL query for user profiles."
-
Thought Generation Strategy:
Propose. You could prompt the LLM: "Given the error message...and the user report..., propose 5 distinct hypotheses for the root cause of the bug, spanning the frontend, backend, and database." This encourages a broad-yet-structured initial exploration. -
State Evaluation Strategy:
Value. For each hypothesis (thought), you could ask the LLM to evaluate its likelihood. The prompt could be: "Evaluate the following hypothesis on a scale of 1-10 based on the provided error logs. A high score means the hypothesis is highly consistent with the evidence. Justify your score." This allows you to rank and prioritize which lead to investigate first. A subsequent thought could be to propose a specific test (e.g., aprintstatement, a unit test) to validate the hypothesis. -
Search Algorithm: DFS (Depth-First Search) would be a natural fit. It would explore the most promising hypothesis (the one with the highest value) first. For instance, it would proceed to generate tests for that hypothesis. If the tests invalidate it, the state value drops, and the algorithm backtracks to explore the next-most-likely hypothesis. This mimics how a human developer systematically investigates and rules out potential causes.
Conclusion
Today, we've elevated our prompting skills from directing an LLM down a single path to architecting a system where the LLM can intelligently navigate a complex problem space. The Tree of Thoughts framework is a powerful example of combining the reasoning capabilities of LLMs with classical AI search algorithms.
Key Takeaways:
- ToT is a Framework, Not Just a Prompt: It orchestrates multiple LLM calls for thought generation and evaluation within a tree search algorithm (like BFS or DFS).
- Four Key Design Choices: Implementing ToT requires defining the thought decomposition, a generation strategy (
sampleorpropose), an evaluation strategy (valueorvote), and a search algorithm. - Flexibility for Different Problems: ToT can be adapted for both highly logical, constrained problems (like math) and open-ended, creative tasks (like writing) by choosing the right combination of strategies.
- Superior Performance on Complex Tasks: By enabling exploration, self-evaluation, and backtracking, ToT can solve problems that are intractable for linear methods like Chain-of-Thought.
Preview of the Next Lesson:
In this module, we've focused on getting the most out of an LLM's internal knowledge and reasoning capabilities. We've structured its thinking process to be more logical, consistent, and now, even exploratory. However, an LLM's knowledge is frozen at the time of its training. What happens when a problem requires information that is too new, too specific, or too proprietary for the model to know?
In our next module, we will begin exploring Retrieval-Augmented Generation (RAG). Our first lesson will be to implement a dense retrieval system using sentence embeddings and a vector database, the foundational step in giving LLMs access to external, up-to-date knowledge before the generation process even begins.