Calories and eating
I photographed my breakfast, then weighed it.
The photo said 600. The photo plus a description said 425. The scales said 514.
One plate, one morning, two estimates 175 calories apart. That is one meal and proves nothing on its own, so the published accuracy studies do the generalising further down. They report something more uncomfortable than a wide spread.
Published 6 September 2026. This is a summary of published research alongside a single self-experiment, not medical or dietary advice.
The plate, and the benchmark
Scrambled egg with spinach and tomato on gluten-free toast, under a generous pile of smoked salmon. Everything meaningful went on the scales.
| Ingredient | Weighed | Contributes |
|---|---|---|
| Egg whites | 185 g | 96 kcal |
| Egg yolks | 2, about 34 g | 109 kcal |
| Smoked salmon | 80 g | 94 kcal |
| Gluten-free bread | 70 g | 175 kcal |
| Spinach | 80 g | 18 kcal |
| Tomato | 120 g | 22 kcal |
| Pepper, chives, lemon | Negligible | 0 kcal |
| Weighed total | 514 kcal |
The plate is written up as a recipe, so the weights below can be reproduced rather than taken on trust.
That comes to 514 kcal, 46 g of protein, 15 g of fat and 43 g of carbohydrate, recomputed here from the weights rather than taken on trust. The spinach was wilted in a dry pan, so there is no cooking fat hiding in the total.
Estimate one: the photograph alone
From the picture and nothing else, the estimate was 600 kcal, with 47 g of protein, 31 g of fat and 32 g of carbohydrate. Against the weighed 514, that is 86 kcal high, or about 17%.
The foods were all identified correctly. Recognition was never the problem. What the photograph could not supply was how many eggs, whether they had yolks in them, how much the salmon weighed, how dense the bread was, and whether anything had been cooked in oil. The scramble was read as roughly three whole eggs, and that single assumption is most of the 31 g of fat.
Estimate two: adding information made it worse
Then one sentence was added: the eggs were mainly whites rather than whole eggs. The estimate fell to 425 kcal, 41 g of protein, 13 g of fat, 33 g of carbohydrate.
The sentence was true and it was incomplete. The plate did contain two yolks, and dropping them took roughly 110 kcal off a meal that had them in it. So the estimate travelled from 17% too high, through the correct answer, to 17% too low.
This is the part worth keeping. More information did not produce a better answer. It produced a confidently different one, and the correction was larger than the error it was fixing.
What happens when somebody does this properly
One breakfast is an anecdote. Fortunately this has been tested at a scale that is not.
Fridolfsson and colleagues photographed 52 standardised items and meals in three portion sizes, gave three models identical prompts with cutlery and plates visible for scale, and compared them against direct weighing. ChatGPT-4o and Claude 3.5 Sonnet came in at a mean absolute percentage error of about 36% for energy, with weight estimation similar. Gemini 1.5 Pro ranged from 64% to 110%.
A second study across three datasets, including 1,463 US meal photographs, found the best-performing model missed by an average of 117 kcal per meal, and that packaged products scored better than meals on a plate. It also found that how you ask matters, but only when the model is capable enough for the asking to help.
So the single breakfast on this page was, if anything, a good day. Seventeen per cent is comfortably better than 36%.
The finding that actually matters
Buried in the same paper is the thing worth changing your behaviour over, and it is not the size of the error.
All three models systematically underestimated, and the underestimate grew with portion size. The bias slopes ran from minus 0.23 to minus 0.50. The bigger the plate, the further under it read.
An error that is random averages out. Log a week of meals with a coin-flip error and the week is roughly right. An error that points the same way every time does not average out, it accumulates, and one that grows with the size of the meal accumulates fastest on exactly the days you most wanted to catch.
That is the mechanism by which a photo-based estimate could quietly remove a deficit rather than track it. Not by being wildly wrong on any one plate, but by being slightly and consistently generous about the big ones.
Which is also the argument for judging the plan on the scale rather than on the app. A weekly average of your weight will catch a systematic logging error that the log itself, by definition, cannot.
But hold on: how good was the old way?
Here is where the obvious conclusion gets complicated, and the authors of that study say it themselves. Their verdict on ChatGPT and Claude was that both achieved accuracy comparable with traditional self-reported dietary assessment methods, without the user burden.
And traditional self-report is worse than its reputation. Pooling five validation studies against doubly labelled water, adults under-reported their energy intake by an average of 28% on a food frequency questionnaire and 15% on a single 24-hour recall, without anybody lying. That figure has its own page, where it sits alongside the equally large errors on the expenditure side.
So the honest comparison is not photograph against truth. It is photograph against a hand-typed diary that was already 15 to 28% under. Both under-report, both under-report more as the meal gets bigger, and one of them takes three seconds.
The benchmark is not exact either
This page has been calling 514 kcal the truth. It is not, quite, and the gluten-free toast shows why.
Seventy grams of it was weighed exactly. But gluten-free bread runs anywhere from roughly 200 to 300 kcal per 100 g depending on the brand, and both ends of that are real products on real shelves. Hold every other weight constant and change only which loaf it was:
- Bread at 200 kcal per 100 g479 kcal
- Bread at 250 kcal per 100 g514 kcal
- Bread at 300 kcal per 100 g549 kcal
A 70 kcal spread on a perfectly weighed meal, from one ingredient, without anybody making a mistake. Weighing answers the question of how much food is there. It does not answer how calorie-dense this particular version of it is, and on processed items that second question can be worth more than the first.
What a photograph genuinely cannot see
- Weight. Eighty grams of smoked salmon and 120 grams look much the same depending on how the slices are folded. Volume estimation from a single image is the open problem in this field, not food recognition.
- What is inside a scramble. Four egg whites, or 185 g of whites with two yolks, or three whole eggs, all occupy about the same space on a plate. On this breakfast that ambiguity alone was worth 110 kcal and 18 g of fat.
- Cooking fat. A tablespoon of olive oil is roughly 120 kcal and it leaves no trace once it is absorbed into eggs. Here there was none, which is itself invisible.
- Which brand it was. The bread above. Also sauces, dressings and anything processed, where two visually identical portions can differ by half again.
The fix is a range and two questions
A tracker that reports 537 kcal from a photograph is claiming a precision the image cannot support. Three significant figures from a picture of a bowl is a design decision, and it is the wrong one, because it hides the uncertainty rather than handing it to the person who could resolve it.
The more honest version reports a band, names its own weakest assumption, and then asks about it. For this breakfast, two questions would have done nearly all the work available:
- Were those whole eggs or mostly whites?
- Was any oil or butter used to cook it?
Neither asks how many grams of tomato there were, because the tomato is 22 kcal and could be wrong by half without mattering. The eggs, the bread and the salmon carry 74% of this plate between them. Those are the three worth weighing, and the rest is rounding.
Which is the practical answer, and it is not photograph or scales. Weigh the two or three things carrying most of the calories, let the picture handle the vegetables, and read the result as a range rather than a verdict.
And check it against something the log cannot fake. If the weekly average is not moving the way the numbers say it should, the numbers are wrong, and after this experiment the most likely direction for them to be wrong in is not the flattering one.
Fuel8 is on the App Store.
A food diary for iPhone. It tracks protein, carbs and fat alongside calories, shows a weight trend rather than a daily verdict, never adds exercise calories back, and has no streaks to break.
Download on the App StoreOr leave your name and email for new recipes and what we add next. No more than that.
This form is the website, not the app. Your name and email go to Grow More Services L.L.C-FZ, the company that builds Fuel8, and are stored in Mailchimp so we can email you about new recipes and what we add next. This list is separate from your optional Fuel8 account. Unsubscribe in one click, any time.
Common questions
How accurate is AI calorie counting from a photo?
Across 52 standardised food photographs, ChatGPT-4o and Claude 3.5 Sonnet both estimated energy with a mean absolute percentage error of about 36%, while Gemini 1.5 Pro ranged from 64% to 110% (Fridolfsson et al., 2025). On a separate benchmark of 1,463 US meal photographs, the best model missed by an average of 117 kcal per meal (Nakagawa & Yamamoto, 2026). The authors of the first study concluded that this is comparable with traditional self-reported dietary assessment, which is a lower bar than most people assume.
Does a photo estimate run high or low?
Low, and increasingly so as the plate gets bigger. All three models tested by Fridolfsson and colleagues showed systematic underestimation, with bias slopes between -0.23 and -0.50, meaning the error grew with portion size. That direction matters more than the size of the error: an estimate that is randomly wrong averages out over a week, and one that is wrong in the same direction every time does not.
Does telling the AI more about the meal make it more accurate?
Only if what you tell it is complete. In the single breakfast on this page, adding the detail that the eggs were mainly whites moved the estimate from 600 kcal to 425 against a weighed 514, so it went from 17% too high to 17% too low. The description left out two yolks. Prompt design does matter in the published work, but Nakagawa and Yamamoto found it only helps when paired with a capable enough model.
What can a photograph not see?
Weight, and anything that was absorbed. Eighty grams of smoked salmon and 120 grams look similar depending on how the slices are folded. A scramble made from four egg whites, from 185 g of whites plus two yolks, or from three whole eggs occupies the same space on a plate with materially different fat. And a tablespoon of oil is about 120 kcal that leaves no visual trace once it is in the eggs.
Is weighing food actually exact?
Closer, but not exact. The weighed benchmark for this breakfast comes to 514 kcal using standard composition values. Change only the gluten-free bread, from 200 to 300 kcal per 100 g, both of which are real products, and the same weighed meal comes to anywhere between 479 and 549. Weighing removes the portion-size question and leaves the food-density one.
So should I photograph my food or weigh it?
It depends what the number is for. If it is for awareness, a photo estimate is in the right region and takes seconds. If you are running a deliberate deficit and the weekly rate is your instrument, a systematic underestimate that grows with portion size is the specific failure that removes a deficit without telling you. The useful hybrid is to weigh the handful of things that carry most of the calories, which for this plate was the eggs, the bread and the salmon, and to photograph the rest.
More in this series: calories eaten versus calories burnt, weigh daily and believe the weekly average, and what a restaurant meal really costs.
Sources
- Fridolfsson J, Sjöberg E, Thiwång M, Pettersson S. Performance evaluation of 3 large language models for nutritional content estimation from food images. Current Developments in Nutrition, 2025. 52 standardised photographs, three portion sizes, weighed reference.
- Nakagawa S, Yamamoto A. Prompt engineering and model selection for LLM-based nutritional estimation from food images: a multi-dataset investigation. Nutrients, 2026. Three datasets, 691 and 1,463 meal photographs and about 1,000 packaged products.
- Tay W, Kaur B, Quek R, Lim J, Henry CJ. Current developments in digital quantitative volume estimation for the optimisation of dietary assessment. Nutrients, 2020;12:1167.
- Freedman LS et al. Pooled results from 5 validation studies of dietary self-report instruments using recovery biomarkers for energy and protein intake. American Journal of Epidemiology, 2014;180:172–88. Doubly labelled water reference.
The breakfast is a single meal estimated twice, which is an anecdote and is treated as one: the general claims on this page come from the cited studies, not from it. Its weighed total was recomputed from the stated weights using standard composition values and came to 514 kcal against the 515 recorded on the day. Estimates were produced in ordinary conversation rather than under a controlled protocol, and different models on different days give different answers. This page summarises published research and is not medical or dietary advice.