Jobs and CronJobs¶
What You'll Learn¶
- Why a Deployment is the wrong tool for run-to-completion work, and what a Job does instead
- The Job fields that control how much work runs, how much runs at once, and how many failures are tolerated
- CronJob scheduling syntax and the settings that prevent overlapping or pileup runs
Why This Matters¶
A Deployment's entire model is "keep N replicas running forever" — feed it a container that's supposed to run once and exit, and it will restart that container forever, which is exactly wrong for batch work. Jobs and CronJobs exist because "run this to completion" and "run this on a schedule" are genuinely different reconciliation problems from "keep this running."
Job: Run to Completion¶
apiVersion: batch/v1
kind: Job
metadata:
name: db-migration
spec:
completions: 1
parallelism: 1
backoffLimit: 3
activeDeadlineSeconds: 900 # give up after 15 minutes, whatever the retries
ttlSecondsAfterFinished: 86400 # delete the Job and its Pods a day after it finishes
template:
spec:
restartPolicy: OnFailure
containers:
- name: migrate
image: myapp-migrate:1.4.2
command: ["./migrate.sh"]
| Field | Controls |
|---|---|
completions |
How many successful Pod completions satisfy the Job overall — 1 for a single task, higher for a fixed batch of independent units of work |
parallelism |
How many Pods can run at once while working toward completions |
backoffLimit |
How many times a failing Pod is retried before the Job itself is marked Failed (with an exponential backoff between retries) |
restartPolicy |
Must be OnFailure or Never for a Job's Pod template — Always (the Deployment default) isn't valid here, since it would fight the whole "run to completion" model |
activeDeadlineSeconds |
A hard wall-clock limit for the whole Job; a hung migration fails instead of running forever |
ttlSecondsAfterFinished |
Automatic cleanup of finished Jobs and their Pods, so they don't pile up |
kubectl create job one-off-task --image=busybox:1.36 -- sh -c "echo done"
kubectl get jobs
kubectl logs job/one-off-task
A Job with completions: 5, parallelism: 2 runs at most 2 Pods concurrently until 5 have succeeded in total — useful for a fixed amount of independent, parallelizable work like processing a batch of files.
CronJob: Jobs on a Schedule¶
apiVersion: batch/v1
kind: CronJob
metadata:
name: nightly-cleanup
spec:
schedule: "0 2 * * *"
timeZone: "Europe/London" # without this, the schedule uses the controller's time zone (usually UTC)
concurrencyPolicy: Forbid
startingDeadlineSeconds: 300
successfulJobsHistoryLimit: 3
failedJobsHistoryLimit: 1
jobTemplate:
spec:
backoffLimit: 2
template:
spec:
restartPolicy: OnFailure
containers:
- name: cleanup
image: myapp-cleanup:1.2.0
command: ["./cleanup.sh"]
A CronJob is a template that creates a new Job object on the schedule you give it — schedule uses standard five-field cron syntax (minute, hour, day-of-month, month, day-of-week; 0 2 * * * means 2 a.m. daily).
timeZone takes an IANA name such as Asia/Kolkata or America/New_York. Without it, "2 a.m." means 2 a.m. in the kube-controller-manager's time zone, which on managed clusters is almost always UTC — a classic reason a "nightly" job runs in the middle of someone's working day. With a real time zone, daylight-saving shifts are handled for you, but a job scheduled inside the skipped or repeated hour may run zero or two times on those days.
| Field | Controls |
|---|---|
concurrencyPolicy |
What happens if the previous scheduled run is still going when the next one is due — Allow (default, runs both), Forbid (skip the new run), Replace (kill the old run, start the new one) |
startingDeadlineSeconds |
How late a missed run can start and still be considered on-time — past this, the run is skipped and counted as missed, rather than run late |
successfulJobsHistoryLimit / failedJobsHistoryLimit |
How many completed Job objects to keep around for inspection before garbage collection |
kubectl get cronjobs
kubectl get jobs --watch # watch new Jobs appear as the schedule fires
kubectl create job --from=cronjob/nightly-cleanup manual-run-now # trigger one run immediately, outside the schedule
When a Job Fails: BackoffLimitExceeded¶
$ kubectl describe job db-migration
...
Conditions:
Type Status Reason
---- ------ ------
Failed True BackoffLimitExceeded
The Job's Pods failed more times than backoffLimit allows (6 if you don't set it), so the Job controller stopped retrying and marked the Job Failed. The Job object is fine. The question is why its Pods failed:
kubectl get pods -l job-name=db-migration
kubectl logs job/db-migration # logs from one of the Job's Pods
kubectl describe pod <failed-pod> # exit code, OOMKilled, image pull errors
- With
restartPolicy: Never, every failure leaves a failed Pod behind to inspect. WithOnFailure, the container restarts inside the same Pod, so checkkubectl logs --previous. - A Job stopped by
activeDeadlineSecondsshows the reasonDeadlineExceededinstead. - A failed Job doesn't restart. Fix the cause, then delete and recreate the Job. For a CronJob, trigger a fresh run with
kubectl create job --from=cronjob/<name>.
Use Cases¶
- Job: database migrations run once per deploy, one-off data backfills, batch image/video processing split across parallel workers, CI-triggered test runs.
- CronJob: nightly backups, periodic cache warming, scheduled report generation, certificate renewal checks, cleaning up expired records.
Common Mistakes¶
- Using a Deployment for a task meant to run once and exit — it will restart the container in a loop forever, since a Deployment has no concept of "done."
- Setting
concurrencyPolicy: Allow(the default) for a job that isn't safe to run twice concurrently — e.g., two overlapping backup jobs writing to the same location. - Forgetting
restartPolicy: OnFailureorNeveron a Job's Pod template — leaving the field unset does not default to something Job-compatible in every context, andAlwaysis rejected. - Not setting
successfulJobsHistoryLimit/failedJobsHistoryLimit(for CronJobs) orttlSecondsAfterFinished(for one-off Jobs), letting finished Job and Pod objects accumulate. - Leaving out
timeZoneand scheduling a CronJob in local time when the cluster runs on UTC. - A retrying Job that isn't idempotent: with
backoffLimit: 3, a migration that half-applied before crashing runs up to three more times. Make the work safe to repeat, or usebackoffLimit: 0and handle failure by hand.
Interview Questions¶
- Why can't you just use a Deployment for a batch task, mechanically?
- What's the difference between
completionsandparallelismon a Job? - What does
concurrencyPolicy: Forbidprotect against on a CronJob, and when wouldReplacebe more appropriate instead?
See Interview Prep for full answers.
Next¶
This closes out Core Concepts. Continue to Workloads and Scheduling for DaemonSets, StatefulSets, deployment strategies, autoscaling, and scheduling controls like affinity and taints/tolerations.