Skip to main content
Create your own
Lesson illustration

Handling and Retrying Failed Jobs

Hello! Welcome back to our course on mastering Laravel.

In our last lesson, we got our background processing system up and running by learning how to start and manage queue workers. You now know that a dispatched job sits in a queue until a worker process picks it up and executes it. But in the real world, things don't always go according to plan. A job might fail due to a bug in your code, a network timeout when calling an external API, or invalid data. A robust application must be able to handle these failures gracefully.

This lesson is dedicated to that very challenge, addressing the learning outcome: Implement strategies for handling and retrying failed jobs. We will explore how Laravel helps you build resilient systems by automatically retrying failed jobs, allowing you to inspect them, re-run them manually, and execute cleanup code when a job fails permanently.

1. The Anatomy of a Failed Job

When a job's handle() method throws an unhandled exception, the worker catches it. By default, the job is considered "failed". But Laravel doesn't just discard it. Instead, it logs the failure for you to inspect later.

To store this information, Laravel uses a dedicated database table: failed_jobs. If you followed the Laravel installation, you should already have a migration for this table. If not, you can create one with php artisan make:queue-failed-table.

This table captures crucial information about the failure:

  • uuid: A unique identifier for the failed job instance.
  • connection and queue: Which queue connection and queue the job was on.
  • payload: The full serialized job object, including all its data.
  • exception: The full stack trace of the exception that caused the failure.
  • failed_at: The timestamp of the failure.
Laravel Failed Jobs Table in PostgreSQL
This image shows what a record in the `failed_jobs` table looks like. You can clearly see the exception message and other details, which are invaluable for debugging.

Let's see a live demonstration of a job failing and how to view it.

Laravel Queues Lesson 2 — Failed jobs: listing, retrying and handling them

This video from Mateus Guimarães will walk you through purposefully failing a job to see how Laravel records it in the failed_jobs table and how you can list these jobs using an Artisan command.

Watch from 01:28 to 02:32. Pay attention to how the failed job is recorded in the database and the output of the php artisan queue:failed command.

As the video shows, the php artisan queue:failed command provides a convenient way to inspect all failed jobs directly from your terminal, without needing to query the database manually.

2. Controlling Retry Attempts and Backoff

Simply logging a failure is useful, but often a failure is temporary (like a brief network glitch). In these cases, we want the job to be retried automatically.

Retry Attempts

In the previous lesson, you saw the --tries option for the queue:work command. While useful, a more common and granular approach is to define the number of attempts directly on the job class itself. This is done by adding a public $tries property.

Backoff Strategy

Retrying immediately isn't always the best strategy. If you're calling an external API that is temporarily down, hammering it with immediate retries won't help. It's better to wait a bit before trying again. This waiting period is called a backoff.

You can define a simple backoff by setting the public $backoff property on your job class to the number of seconds you want to wait between attempts.

For more complex strategies, you can define a backoff() method that returns an array of delays. This allows you to implement an exponential backoff, where the delay increases after each failed attempt.

/**
 * The number of times the job may be attempted.
 *
 * @var int
 */
public $tries = 5;

/**
 * Calculate the number of seconds to wait before retrying the job.
 *
 * @return array<int, int>
 */
public function backoff(): array
{
    return [1, 5, 10];
}

In this example:

  • The job will be attempted a total of 5 times.
  • If the 1st attempt fails, it will wait 1 second before the 2nd attempt.
  • If the 2nd attempt fails, it will wait 5 seconds before the 3rd attempt.
  • If the 3rd attempt (and any subsequent attempts) fail, it will wait 10 seconds.

Let's see how setting the number of tries works in practice.

Laravel Queues Lesson 2 — Failed jobs: listing, retrying and handling them

This clip demonstrates how to set the max attempts both on the worker command and, more importantly, within the job class itself.

Watch from 03:11 to 06:06. Notice the difference in the worker's output: the yellow 'failed' messages are temporary attempts, while the final red 'failed' message indicates the job has exhausted its retries and is being moved to the failed_jobs table.

To dive deeper into all the configuration options for retries and backoffs, the official documentation is the best place to go.

Queues - Laravel Documentation

The Laravel documentation provides a comprehensive overview of configuring job attempts and backoff strategies. It's your reference for all available options.

Read the section 'Dealing With Failed Jobs'. Focus on the sub-sections about the --tries switch, the --backoff option, and how to define $tries, $backoff, and the backoff() method directly on your job class.

3. Managing and Retrying Failed Jobs Manually

Once a job has exhausted all its automatic retry attempts, it lands in the failed_jobs table and stays there. Now it's up to you, the developer, to decide what to do. Perhaps you've fixed the underlying bug or the external service is back online. You can now manually retry the job.

Laravel provides simple Artisan commands for this:

  • php artisan queue:retry [UUID]: Retries a single job using its UUID (which you get from queue:failed).
  • php artisan queue:retry all: Retries all jobs in the failed_jobs table.
  • php artisan queue:forget [UUID]: Deletes a single failed job from the table.
  • php artisan queue:flush: Deletes all jobs from the failed_jobs table.

Let's watch a demonstration of these commands.

Laravel Queues Lesson 2 — Failed jobs: listing, retrying and handling them

This final clip from the same video shows how to use the job's UUID to retry it after fixing the error, and how to retry all failed jobs at once.

Watch from 02:26 to 03:16. See how the job is removed from the failed_jobs table once it is successfully retried.

4. Performing Actions on Final Failure

What if you need to perform a specific action when a job fails for good? For example, you might want to send a notification to a Slack channel, alert a developer, or reverse a partial database transaction.

Laravel allows you to do this by defining a failed() method in your job class. This method will be automatically called when the job has exhausted all its retry attempts and is about to be written to the failed_jobs table. It receives the Throwable exception that caused the final failure as an argument.

use Throwable;

/**
 * Handle a job failure.
 */
public function failed(?Throwable $exception): void
{
    // Send a notification to the development team.
    // Log the error to an external monitoring service.
    // Revert a database state change if necessary.
}

This hook is incredibly powerful for building self-healing and observable systems. Let's see it in action.

Laravel Queues Lesson 2 — Failed jobs: listing, retrying and handling them

This segment demonstrates how the failed() method is invoked only after all retry attempts are exhausted, allowing you to run custom logic at the point of permanent failure.

Watch from 06:06 to 07:35. Notice that the failed() method's code only runs on the final, red 'failed' attempt, not on the intermediate yellow ones.

5. Best Practices: Designing for Failure

Implementing retries is a mechanical process. The real art is designing your jobs so that retries are safe. The key concept here is idempotency.

An idempotent operation is one that can be performed multiple times with the same result as performing it once.

  • Not idempotent: Sending an email. Retrying this job will send multiple emails.
  • Idempotent: Updating a user's status to processed. Retrying this job will just set the status to processed again, which has no negative side effect.

When building jobs that are not naturally idempotent (like charging a credit card), you need to build in checks to prevent duplicate actions. For example, before charging the card, check if an order has already been marked as paid.

Queues That Don't Fail

This blog post provides excellent high-level strategies for designing robust queue systems, with a great section on idempotency.

Read the sections 'Idempotency and safe side effects' and 'What to do with failed jobs'. These sections will give you a strong conceptual foundation for why these retry mechanics matter and how to use them effectively.

For a comprehensive view of failure management, especially in production, tools like Laravel Horizon provide a beautiful dashboard to monitor your queues and failed jobs, allowing you to inspect and retry them with a click of a button.

Failed Jobs Management Interface
This is a screenshot of the Laravel Horizon dashboard, which gives you a user-friendly interface for viewing and retrying failed jobs, complete with stack traces and other details.

Conclusion

You have now learned the essential strategies for building resilient, production-ready background job processing systems in Laravel. You can control how and when jobs are retried, and what to do when they ultimately fail.

Key Takeaways:

  • Jobs that fail are stored in the failed_jobs database table for inspection.
  • You can view failed jobs with php artisan queue:failed.
  • Control automatic retries using the $tries property and $backoff property/method in your job class.
  • Manually re-run failed jobs using php artisan queue:retry by ID or all.
  • Execute cleanup logic or send notifications upon permanent failure using the failed() method.
  • Designing idempotent jobs is crucial for making retries safe and predictable.

Up Next:

Now that we have a robust system for handling both successful and failed jobs, we'll apply this knowledge to a very common and practical use case. In the next lesson, we will focus on how to offload long-running tasks like sending emails or notifications to a queue, which is a primary reason for using queues in the first place.

Can't find a good explanation? Sign up and we'll make it for you

Sign up