Ceph capacity and failure-domain planner

Version 1.1.0

A cluster's raw disk total doesn't tell you how much data Ceph can protect. Enter the drives, where they are grouped, and the replication or erasure-coding profile to estimate usable capacity. The planner also checks whether enough separate failure domains remain when the largest one is removed.

Protection mode Required
Replica size Required
Failure domain Required
Planning reserve
%

OSD batches

0 / 32 batches · 0 / 256 OSDs

How replication and erasure coding change Ceph capacity

Replication stores complete copies. With a replica size of three, the capacity model uses three raw bytes for each byte of data. Erasure coding splits data into k data chunks and adds m recovery chunks, so the data fraction is k / (k + m).

For a hypothetical 120 TB raw inventory, these ratios give the following estimates before operating reserve or live-cluster overhead:

ProfileData fractionModeled usable capacity
Three replicasOne third40 TB
Erasure coding, four data and two recovery chunksTwo thirds80 TB

The larger number is not a recommendation. Ceph's erasure-coding documentation explains the workload and recovery tradeoffs that a capacity ratio cannot measure.

Why the failure-domain labels matter

A failure domain is a group that can disappear together, such as a host or rack. Six drives in one host are not six independent hosts. Enter batches under the same label when they belong to the same modeled domain; the tool groups their capacities.

The selected model needs as many separate domains as replicas, or k + m domains for erasure coding. Passing that count checks an input assumption, not the placement of real data. Ceph's CRUSH rules determine where copies or chunks actually go.

The largest-domain scenario removes the biggest entered group and recalculates with the same ratio. It shows the effect on this inventory model. It doesn't simulate recovery, prove that enough copies survive, or guarantee that writes can continue.

What this model leaves out

This calculation does not inspect live CRUSH buckets, pool rules, PG counts, min_size, object overhead, BlueStore overhead, recovery, IOPS, or actual ceph df output. It also does not prove durability after a domain failure. The loss scenario is a planning model that removes the largest entered domain and keeps the selected ratio.

Reserve is a planning policy, not a Ceph nearfull threshold. Read the Ceph pool configuration reference when live pool settings affect the question.

Verify the plan against a real cluster

Compare the entered batches with the real OSD and CRUSH hierarchy before acting on the estimate. Check the pool rule and profile as well as the capacity view.

  • ceph osd tree and ceph osd crush tree show the actual hierarchy.
  • ceph osd pool get <pool> size and min_size show pool policy.
  • ceph osd erasure-code-profile get <profile> shows live EC settings.
  • ceph df shows the cluster view after deployment.

See the Ceph CRUSH map documentation and EC profile documentation. For adjacent storage planning, return to storage tools or compare a ZFS capacity calculator.

Sharing keeps the calculation explicit

Calculation stays in your browser, and the tool makes no calculation request. Share is explicit: it puts normalized inputs in a versioned URL. Anyone with that URL can read the entered inventory labels and capacities.