Benford’s law says that in many real-world datasets the leading digit is 1 about 30.1% of the time, 2 about 17.6%, and 9 only 4.6% (the share for digit d is log10(1 + 1/d)). Invented or manipulated numbers often break this pattern, which is why auditors use it to flag data for review. Paste your numbers to get the first-digit distribution, a chart, a chi-square test and the mean absolute deviation (MAD) verdict auditors use. Example: the powers of 2 from 21 to 2300 conform closely (MAD = 0.0013), while 500 random three-digit numbers do not (MAD = 0.062). A mismatch is a reason to look closer, not proof of fraud.

Paste a spreadsheet column. Currency symbols and thousands separators (1,234.56) are fine; use a dot for decimals. Zeros and blanks are ignored; negative numbers use their absolute value.

What Benford’s law is

In many collections of numbers that span several orders of magnitude — invoice amounts, populations, river lengths, stock volumes — small leading digits are far more common than large ones. The expected share of numbers starting with digit d is log10(1 + 1/d): 30.1% for 1, 17.6% for 2, 12.5% for 3, down to 4.6% for 9. The pattern appears whenever numbers grow multiplicatively or combine several scales, which is why it shows up in so much real data.

Using it to spot suspicious data

People who invent numbers tend to spread first digits too evenly, or overuse particular digits. Forensic accountants (notably Mark Nigrini) use first-digit tests to choose which accounts or transactions to audit. A failed test is a flag for review: there are many innocent reasons for deviations, such as price points, thresholds and caps.

When Benford’s law does not apply

  • Narrow ranges — adult heights or exam scores sit within one order of magnitude.
  • Assigned numbers — phone numbers, ZIP codes, invoice IDs.
  • Built-in limits — prices set at 9.99, expense claims capped just under an approval limit, minimums and maximums.
  • Small samples — with fewer than about 100 numbers, chance variation swamps the pattern.

Chi-square or MAD?

The chi-square test becomes so sensitive with large datasets that trivial deviations look “significant”. Auditors therefore rely on the mean absolute deviation between observed and expected shares, which does not grow with sample size, and use the chi-square test as a secondary signal. Our worked examples show both: the powers of 2 conform closely (MAD 0.0013, χ² 0.06), while 500 uniformly random three-digit numbers clearly do not (MAD 0.062, p < 0.0001).

Common mistakes

  • Treating a failed test as proof of fraud. It only says the data is unusual.
  • Testing data that should not follow the law in the first place (see above).
  • Mixing very different datasets into one test, which can hide or create deviations.

Related

Chi-square testDescriptive statisticsProbability calculatorAll tools

Frequently asked questions

It states that in many naturally occurring datasets the first digit is 1 about 30.1% of the time and 9 only about 4.6% of the time, following log10(1 + 1/d).

It can flag data that deserves a closer look, because fabricated numbers often do not follow the expected digit pattern. A deviation is not proof of fraud, since many legitimate datasets deviate.

At least about 100, and preferably several hundred or more, spanning at least two orders of magnitude.

Computed in your browser; formulas verified against SciPy, statsmodels and reference packages. See the methodology page.