Our worker deployment sat at one replica while forty thousand jobs piled up in Redis. The HPA was healthy, the target was 70% CPU, and actual CPU was 2%. Nothing was wrong — the autoscaler was measuring the wrong number, because a Laravel worker blocked on Redis is asleep whether the queue holds zero jobs or forty thousand.
This walks through replacing that HPA with a KEDA ScaledObject that scales on queue depth, scaling the deployment to zero when the queue is empty, and configuring shutdown so Kubernetes stops killing jobs halfway through.
Understand why the CPU-based HPA cannot see your queue#
A Laravel worker spends almost all of its life inside a blocking Redis call or a sleep(). Neither consumes CPU. Graph worker CPU next to queue depth during a backlog and you get a flat line under a climbing one — the HPA is reading a metric that is genuinely uncorrelated with load, so it correctly concludes nothing needs to happen.
KEDA fixes the metric, not the mechanism. It runs an operator plus a Kubernetes external metrics adapter: you declare a ScaledObject, KEDA creates and owns an ordinary HPA behind it, and feeds that HPA a metric from Redis instead of from cAdvisor. Everything the cluster already knows about HPA behaviour, stabilisation windows and scaling policies still applies. If you are new to tuning the worker side of this, the production guide to scaling Laravel queues covers the process-level decisions that sit underneath the pod-level ones here.
Install KEDA into the cluster#
KEDA installs cluster-wide via Helm and adds the CRDs the rest of this article uses: ScaledObject, ScaledJob, TriggerAuthentication and ClusterTriggerAuthentication. It needs permission to install cluster-scoped resources.
helm repo add kedacore https://kedacore.github.io/charts
helm repo update
helm install keda kedacore/keda \
--namespace keda \
--create-namespace \
--version 2.21.0
# Both pods must be Running before any ScaledObject will work:
# the operator reconciles CRDs, the metrics apiserver serves the HPA.
kubectl get pods -n keda
kubectl get apiservice v1beta1.external.metrics.k8s.io
If v1beta1.external.metrics.k8s.io reports False (FailedDiscoveryCheck), the metrics apiserver is not reachable and every scaler will report a metrics error rather than a wrong number. Fix that before going further.
Find the Redis key Laravel actually writes to#
This step saves more debugging time than everything else combined. Laravel's RedisQueue::getQueue() returns queues:{name}, and the phpredis client then prepends the prefix from config/database.php. That prefix is not stable across skeleton versions — older Laravel apps generate laravel_database_, current ones generate a slug of APP_NAME with hyphens, like my-app-database-. Guessing is how people end up with a scaler that reports zero forever.
Look at the real keyspace instead:
kubectl exec -it deploy/redis -- redis-cli --scan --pattern '*queues*'
# my-app-database-queues:default <- the list. LLEN counts these.
# my-app-database-queues:default:delayed <- sorted set, NOT counted by LLEN
# my-app-database-queues:default:reserved <- sorted set, in-flight jobs
# my-app-database-queues:default:notify <- list used by block_for wakeups
# my-app-database-queues:exports
Two things follow. First, the listName you give KEDA is the full prefixed key — my-app-database-queues:default, not default. Second, LLEN counts only jobs that are ready right now. Delayed jobs and reserved (in-flight) jobs live in sorted sets that the list scaler does not read, so a queue of nothing but retries will show a depth of zero. If your workload is retry-heavy, keep minReplicaCount: 1 so there is always a worker around to migrate the delayed set back onto the list.
One Laravel 13 wrinkle: getQueue() now passes the queue name through resolveQueue(), so an app using queue routing to forward jobs between queues may write to a key that does not match the queue name in your job class. Trust --scan, not the source code.
Write the ScaledObject#
Here is the complete resource, annotated. Give it the same namespace as the worker deployment — KEDA requires the target to be in the ScaledObject's namespace.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: queue-worker-default
namespace: production
spec:
scaleTargetRef:
name: queue-worker-default # the worker Deployment
pollingInterval: 15 # how often KEDA checks Redis while at 0 replicas
cooldownPeriod: 180 # idle seconds before scaling back to 0
minReplicaCount: 0 # scale to zero when the queue is empty
maxReplicaCount: 20
fallback:
failureThreshold: 3 # if Redis is unreachable 3 polls running...
replicas: 2 # ...hold at 2 rather than dropping to 0
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # stop replica thrashing on spiky queues
triggers:
- type: redis
metricType: AverageValue
metadata:
address: redis.production.svc.cluster.local:6379
listName: my-app-database-queues:default
listLength: "25" # jobs per replica — see below
activationListLength: "1" # wake from 0 as soon as 1 job lands
enableTLS: "false"
authenticationRef:
name: keda-redis-auth
listLength is the field everyone misreads. It is not a threshold that scaling starts at — with the default AverageValue metric type, the HPA computes replicas = ceil(queueDepth / listLength). At listLength: "25", a depth of 300 asks for 12 replicas, capped at maxReplicaCount. Set it to roughly the number of jobs one worker clears in the time you are willing to wait.
activationListLength is a separate gate and only governs zero-to-one. Leave it at its default of 0 and KEDA treats any non-zero depth as active; set it to 5 and four jobs will sit in an empty queue indefinitely. For anything a user is waiting on, 1 is the only sane value.
cooldownPeriod applies only to the scale-to-zero transition. Scaling from one to N is pure HPA, which is why the scaleDown stabilisation window above exists separately — they control different halves of the curve.
Decide where scale-to-zero is actually right#
Zero is the feature that pays for the setup, and it is not free. A nightly reporting queue, per-tenant queues on a low-volume plan, an exports queue that fires twice a day — these cost nothing for most of the day and scale up within pollingInterval plus pod start time. On a Laravel image with a warm layer cache that is typically ten to thirty seconds, which nobody notices on a batch job.
They notice it on a password reset email. Anything synchronous-adjacent gets minReplicaCount: 1 and keeps one warm worker. The rule I use: if a human is looking at a spinner that depends on the job, you do not scale that queue to zero.
Configure shutdown before Kubernetes kills a job#
Scale-down is where data goes missing. Kubernetes sends SIGTERM, the Laravel worker finishes the job it is on and exits cleanly, and Kubernetes sends SIGKILL once terminationGracePeriodSeconds elapses — default 30 seconds. A five-minute export is dead at second thirty, mid-write, with no failed-job record.
Set the grace period above your longest job's runtime:
apiVersion: apps/v1
kind: Deployment
metadata:
name: queue-worker-default
namespace: production
spec:
replicas: 1 # KEDA takes ownership of this
template:
spec:
terminationGracePeriodSeconds: 900 # must exceed the longest job
containers:
- name: worker
image: registry.example.com/my-app:latest
command:
- php
- artisan
- queue:work
- redis
- --queue=default
- --tries=3
- --max-time=3600 # recycle the process hourly to cap memory drift
- --max-jobs=1000
- --timeout=600 # per-job limit; keep below the grace period
lifecycle:
preStop:
exec:
# Let endpoint removal settle before SIGTERM reaches PHP.
command: ["/bin/sh", "-c", "sleep 5"]
Three details matter here. Use queue:work, never queue:listen — listen boots a fresh framework per job and makes cold starts dramatically worse. Keep --timeout below terminationGracePeriodSeconds so the worker's own guard fires before the kernel's. And on Laravel 13, --stop-when-empty-for=60 is a neat companion to scale-to-zero: the worker exits on its own after a minute of silence, so the pod is already gone by the time KEDA's cooldown expires.
The max-jobs and max-time options behave exactly as they do outside Kubernetes. What changes is who restarts the process: KEDA and the Deployment controller, not Supervisor. If you are migrating from a VM setup, the Supervisor configuration for Laravel workers is the direct analogue — same intent, one layer down. The pod-level plumbing, if you have not containerised the worker yet, is covered in taking a Laravel Dockerfile to a running pod.
Authenticate to Redis with a TriggerAuthentication#
Do not put the Redis password in the ScaledObject; it is a plain CRD that everyone with namespace read access can see. TriggerAuthentication pulls it from a secret instead.
apiVersion: v1
kind: Secret
metadata:
name: redis-credentials
namespace: production
type: Opaque
stringData:
password: "s3cr3t"
---
apiVersion: keda.sh/v1alpha1
kind: TriggerAuthentication
metadata:
name: keda-redis-auth
namespace: production
spec:
secretTargetRef:
- parameter: password
name: redis-credentials
key: password
For a managed Redis with TLS, set enableTLS: "true" in the trigger metadata. Only reach for unsafeSsl: "true" with a self-signed cert you control — it disables certificate verification entirely. The same connection details apply if you have moved the queue connection to Valkey; the list scaler speaks the same protocol.
Split one scaler per queue#
A single scaler across mixed queues starves the fast one. If exports holds 400 slow jobs and notifications holds 6 quick ones, a combined scaler sizes the fleet for exports and the notifications sit behind them. Give each queue its own Deployment, its own worker command, and its own ScaledObject with its own listLength:
# notifications: fast jobs, stay warm, aggressive ratio
spec:
scaleTargetRef:
name: queue-worker-notifications
minReplicaCount: 1
maxReplicaCount: 30
triggers:
- type: redis
metadata:
listName: my-app-database-queues:notifications
listLength: "10"
---
# exports: slow jobs, scale to zero, one worker per few jobs
spec:
scaleTargetRef:
name: queue-worker-exports
minReplicaCount: 0
cooldownPeriod: 600
maxReplicaCount: 6
triggers:
- type: redis
metadata:
listName: my-app-database-queues:exports
listLength: "3"
Horizon, if you run it, is not a competitor here — it balances processes inside a pod while KEDA manages the number of pods. They coexist, but pin Horizon's maxProcesses to a fixed value per pod rather than letting its own auto-balancing fight KEDA's replica maths.
Watch queue depth against replica count#
KEDA's state is visible through the HPA it creates and through the ScaledObject itself. These three commands answer almost every "why is it not scaling" question.
# The HPA KEDA owns. The TARGETS column shows actual/target queue depth.
kubectl get hpa -n production
# NAME TARGETS MINPODS MAXPODS REPLICAS
# keda-hpa-queue-worker-default 312/25 (avg) 1 20 13
# Ready=True, Active=True means the trigger is firing. Read the events.
kubectl describe scaledobject queue-worker-default -n production
kubectl logs -n keda deploy/keda-operator --tail=50
The single most useful dashboard panel is queue depth plotted against replica count across one load spike. If depth climbs while replicas sit flat, the scaler is broken. If replicas oscillate while depth is steady, cooldownPeriod or the stabilisation window is too short. Pipe the worker logs somewhere you can correlate them — centralised logging with Grafana and Loki makes the "did that job actually finish before the pod went away" question answerable.
Diagnose the five failures that bite in production#
The key name is wrong. The scaler reports 0 forever, never scales, and logs nothing alarming — a wrong key and an empty queue look identical. kubectl describe scaledobject shows Active: False while Redis clearly has jobs. Re-run redis-cli --scan and compare character for character, including the prefix.
An old HPA is still in place. Two controllers writing spec.replicas on one deployment produces replica flapping with no obvious cause. KEDA's admission webhook rejects a ScaledObject whose target already has an unmanaged HPA; if you want KEDA to adopt it, annotate with scaledobject.keda.sh/transfer-hpa-ownership: "true". Otherwise delete the old HPA first.
cooldownPeriod is too short. At 30 seconds on a bursty queue, the deployment scales to zero, a job lands, it scales back, and repeats. Every cycle pays a cold start. Start at 180–300 seconds and shorten only if you can see idle pods costing real money.
The grace period is still 30 seconds. kubectl get events shows the pod moving to Killing and the container exiting with code 137 while the job's log line never reaches "processed". Raise terminationGracePeriodSeconds above the longest job and re-check after anyone adds a slow one.
Scale-to-zero on a retry-heavy queue. Failed jobs go to the :delayed sorted set, which LLEN does not count, so with zero replicas there is nothing running to migrate them back onto the list when they come due. A queue that is mostly retries needs minReplicaCount: 1.
Wrapping Up#
Install KEDA, discover the real Redis key with --scan, and start with one ScaledObject on your least critical queue — minReplicaCount: 0, listLength at roughly one worker's throughput per acceptable wait, and terminationGracePeriodSeconds set above your slowest job. Watch queue depth against replicas for a week before you tune anything.
Once workers scale themselves, the next gap is usually pods not reporting health honestly. Kubernetes probes for Laravel and a container healthcheck for the worker itself stop a wedged worker from counting as capacity.
FAQ#
How do I autoscale Laravel queue workers on Kubernetes?
Install KEDA and create a ScaledObject pointing at your worker Deployment with a redis trigger that monitors the Laravel queue list. KEDA creates and owns an HPA behind the scenes, feeding it queue depth from Redis rather than CPU. You set listLength to the number of jobs one worker should handle, and the HPA divides current queue depth by that figure to decide the replica count.
Why doesn't the Kubernetes HPA scale my queue workers?
Because the default HPA scales on CPU, and a Laravel worker waiting for a job consumes almost none. The process is blocked on a Redis call or sleeping between polls, so CPU stays near idle whether the queue is empty or holds tens of thousands of jobs. The metric is uncorrelated with load, which is why KEDA exists — it supplies queue depth as an external metric so the same HPA machinery can act on something meaningful.
How do I configure the KEDA Redis scaler for a Laravel queue?
Use type: redis with address, listName and listLength. The listName must be the fully prefixed key, such as my-app-database-queues:default — find it with redis-cli --scan --pattern '*queues*' rather than guessing, because the prefix varies by Laravel skeleton version and APP_NAME. Supply the Redis password through a TriggerAuthentication referencing a Kubernetes secret rather than inlining it in the ScaledObject.
Can Laravel queue workers scale to zero?
Yes — set minReplicaCount: 0 and KEDA removes every replica once the queue has been empty for cooldownPeriod seconds. The trade-off is cold start latency on the first job, typically ten to thirty seconds. It suits batch and low-volume queues; for anything a user is waiting on, keep minReplicaCount: 1. Avoid zero on retry-heavy queues, because delayed jobs live in a sorted set that LLEN does not count.
How do I stop Kubernetes killing a worker mid-job?
Set terminationGracePeriodSeconds on the worker Deployment to a value higher than your longest-running job. Kubernetes sends SIGTERM, which Laravel's worker handles by finishing the current job and then exiting, but it sends SIGKILL once the grace period expires — and the default is only 30 seconds. Keep the --timeout option below the grace period so the worker's own guard fires first, and add a short preStop sleep so endpoint removal settles before the signal arrives.
What KEDA queue key does Laravel use with a Redis queue?
Laravel builds the key as the Redis prefix from config/database.php plus queues: plus the queue name, giving something like my-app-database-queues:default. Older skeletons produce laravel_database_queues:default instead. Three sibling keys exist alongside it — :delayed, :reserved and :notify — and the list scaler reads only the base key, so delayed and in-flight jobs are not reflected in the metric.