What the Piotroski F-Score was built to answer
The score comes from a 2000 paper by Joseph Piotroski, who was studying a specific problem: among companies that already look statistically cheap, which ones are cheap because the business is improving and which are cheap because it is deteriorating? His answer was not a valuation model. It was a checklist of nine accounting signals, each reduced to a single binary point, summed into a score from 0 to 9.
That design choice matters for how the Piotroski F-Score should be read. Each criterion is deliberately crude: better or not better, positive or not positive. The aim was robustness rather than precision. A framework that survives being applied to thousands of filings by different people has to be blunt, and this one is.
The nine criteria, in their three groups
Four criteria cover profitability. One point if net income is positive. One if operating cash flow is positive. One if return on assets improved on the prior year. One if operating cash flow exceeds net income, which is the accruals test. A company reporting profit it has not yet collected in cash scores zero here.
Three cover leverage, liquidity and dilution. One point if long-term debt as a share of assets fell. One if the current ratio improved. One if the share count did not increase, so a company funding itself by issuing stock loses this point.
Two cover operating efficiency. One point if gross margin improved on the prior year, and one if asset turnover improved. Every criterion in this set is a comparison against the company's own previous year, not against a peer group or an industry average, which is why the score can be computed without any external data at all.
The accruals test is the interesting oneComparing operating cash flow with net income is the single criterion that most often separates two companies with identical headline profits. Earnings can be recognized before the money arrives; cash flow cannot. When those two numbers diverge persistently, the calculation is telling you something the income statement alone will not.
A worked calculation, point by point
Take a constructed company. Net income 40 (positive, one point). Operating cash flow 55 (positive, one point; and it exceeds net income, one more point). Return on assets 6.2 percent against 5.1 last year (improved, one point). Long-term debt to assets 0.28 against 0.33 (fell, one point). Current ratio 1.9 against 1.7 (improved, one point). Shares outstanding unchanged (one point). Gross margin 41 percent against 42 (worse, no point). Asset turnover 0.71 against 0.68 (improved, one point).
That is eight of nine. The calculation took two annual reports and no judgment calls, which is the whole appeal of the criteria being binary. Any two people working from the same filings should reach the same integer, and if they do not, one of them has read a line item differently, a disagreement that is findable rather than a matter of opinion.
Piotroski's convention was to treat 8 or 9 as strong and 0 to 2 as weak, with the middle inconclusive. Those bands are his, from the original study, and they describe the signal that study measured. They are not a rating scale, and the score itself carries no information about price.
A second example, this time a failing one
The score is more useful when it goes badly, so here is the same exercise on a company scoring 3. Net income 12, positive, one point. Operating cash flow minus 8: no point, and because it is below net income the accruals test fails as well, so two of the four profitability points are already gone. Return on assets 1.1 percent against 2.4 last year: worse, no point. Long-term debt to assets 0.41 against 0.34: it rose, no point. Current ratio 1.1 against 1.4: worse, no point. Shares outstanding 94 million against 81 million: an increase, no point. Gross margin 33 percent against 31: improved, one point. Asset turnover 0.55 against 0.52: improved, one point.
Three of nine, and the three that were earned are the least informative of the set. What the failures say together is more specific than the total: the company reported a profit it did not collect in cash, funded the gap by borrowing and by issuing stock, and ended the year less liquid than it started. Each of those is a line item, and each points at a page of the filing rather than at a judgment.
That is the practical value of a binary checklist. A graded score would have blurred these into a middling number; nine yes-or-no answers leave a trail. The two efficiency points also show why the total alone misleads: margin and turnover both improved, which in isolation reads as a business getting better at its operations while its financing position deteriorated underneath.
What the Piotroski F-Score does not tell you
It does not tell you a company is undervalued. The original work applied the criteria within a universe already filtered for low price-to-book, so the score was a second stage rather than the whole screen. Applied to an expensive company, a score of 9 says the fundamentals improved on nine specific measures and says nothing whatsoever about what you would be paying for them.
It also has no view on the future. Every one of the nine criteria compares this year with last year, which makes the score a description of a change that has already happened. And because the inputs are annual figures, it is stale by construction: the calculation reflects a filing, not the business as it stands today.
Finally, the criteria assume a conventional balance sheet. Financial companies, early-stage businesses with negative earnings and firms with unusual accounting can produce scores that are arithmetically correct and economically meaningless. The screening value of the F-Score is real, and it is confined to the kind of company the criteria were designed around.
Where the criteria came from, and why that matters
The nine tests were not chosen for elegance. Piotroski was working on a specific empirical problem. Value portfolios contain a large number of genuinely deteriorating businesses, and the average return of the group hides that, so he needed signals that were available for every company, computable without judgment, and robust to being applied mechanically across thousands of filings. Accounting signals that any competent reader can verify satisfy all three; anything requiring an estimate satisfies none.
That origin explains both the strength and the boundary. Inside a value universe the criteria separate improving from deteriorating businesses, which is what they were fitted to do. Outside it they are simply nine facts about a company's last two years, and treating those facts as a quality ranking asks the score for something it was never tested on.
Using it as a screening step rather than an answer
The practical way to use this in screening is as a filter applied after a valuation criterion, not instead of one, and as a prompt to look at particular line items rather than as a verdict. A company scoring 3 is telling you which specific tests it failed, and those are the pages of the filing worth opening.
The worksheet on this site scores the nine criteria one at a time and shows which points were awarded, precisely so the output is a list of findings rather than a single number. That is the honest form for a composite score: inspectable, reproducible, and clear about what it left out.
