Stronger by Math

Training decisions, run through the arithmetic

A note on geometry

Why Every Food Camera Fails on Bowls, and None of Them Say So

From directly above, a bowl with two inches of rice looks identical to one with four. This is optics, not software, and it will not be fixed.

By Caleb Ostrowski·Published ·287words

Photograph a bowl of rice from directly overhead. Add two inches of rice. Photograph it again from the same position.

The images are nearly identical. The calorie content is roughly double.

This is not a bug

Depth is not recoverable from a single monocular photograph without a reference in frame. That is a statement about optics, not about model quality, and it applies to every app in this category regardless of the compute behind it.

We tested six scanners against the same deep bowl. All six overestimated, most by around a third, and they did so consistently. The consistency is the giveaway — this is a systematic error with a physical cause, not noise.

Why it matters more than it sounds

Overhead is how everybody photographs a bowl. It is the natural angle: you are standing over the table holding a phone.

So the most common food-photography posture produces the least tractable geometry, and a large share of home meals are served in bowls. This is not an edge case.

What separates the apps

Given that they all fail the same way, the useful question is what happens after the failure.

An app with a verified catalogue underneath lets you adjust the portion against a real record in two taps. An app that is only a camera drops you into a thin generic list where the correction is often worse than the error.

That question — not the headline accuracy percentage — is what actually separates these tools in daily use.

The habit worth adopting

Photograph bowls at an angle, and put a utensil or your hand in frame for scale. It is not a fix. It measurably helps, and it costs nothing.

Questions we get asked

Why do calorie apps overestimate food in bowls?

Because the photograph does not contain the information needed. From overhead, a bowl filled to two inches and the same bowl filled to four produce nearly the same image — same outline, same visible surface, same colours. The only difference is depth, which a single overhead frame does not encode, so the system falls back on a learned prior about how full bowls usually are. Those priors skew full, which produces a systematic upward bias rather than random noise.

Can I make photo logging more accurate?

Two things help materially and cost nothing. Photograph at an angle rather than straight down, which supplies the fill height the overhead frame omits. And put something of known size in frame. Neither is a fix, and both measurably help.

Caleb Ostrowski

Editor · Stronger by Math

Coaches lifters and writes the arithmetic down. Nine years of client logs, most of them unglamorous. Buys every app on this site at retail and cancels most of them. No affiliate links anywhere on this desk.

More from the desk