Just out todayAI agents & MCP: What a 49.1% attack rate does not tell youCybersecurity: The MCP scanner number that should worry youSpace tech: Sell insurers a one-page orbit crowding score

Data/Agritech/Ethiopia

97.3% in a paper is not 97.3% in a field

An offline crop disease model scored 97.3 percent on cactus-fig photos taken at university test plots. No farmer has ever used it, and the paper describes that figure two different ways.

The short answer

97.3 percent is how often a 9.3 MB model sorted a cactus-fig photo into the right group, on a dataset of 3,587 field photos built at Mekelle University in Ethiopia. We score the preprint 4 out of 10. It is a score on photos the team shot at test plots, and no farmer trial is reported.

Grade 5 reading level5 min read

The number

97.3 percent. That is how often the best of three small models put a photo of a cactus pad in the right group.

The crop is cactus-fig, known as beles in Tigray, Ethiopia. It feeds people in the hungry months before harvest. The model is 9.3 MB, so it fits inside a phone app that runs with no signal.

Our full write-up is here: Crop disease model hits 97.3% and runs offline on old phones.

Where it comes from

The paper comes from Mekelle University and its Mekelle Institute of Technology in Ethiopia. It is a preprint from December 2025, and a shorter version was presented at a regional conference in Mekelle in February 2025.

The team took 3,587 field photos of cactus pads. 1,500 are marked affected, 1,500 healthy, and 587 show no cactus at all. They were shot outdoors at test plots at Mekelle University and Adigrat University, with dust, shadow and glare left in. The test set held 1,195 photos the models never saw.

They trained three models on the same photos. MobileViT-XS scored 97.3 percent, EfficientNet-Lite1 scored 90.7 percent, and a small custom network scored 89.5 percent.

We score it 4 out of 10. The strong point is openness. The code and the photos are public on GitHub and Kaggle, so anybody can check it. It loses points for resting on one dataset, from one university, with no field trial.

What it does not mean

It does not mean the model is right 97 times in 100 in a farmer’s hand. It was right 97 times in 100 on photos the team took at university test plots. A farmer holds a phone differently. The pad is dusty in a different way. Evening light is not noon light. Each of those moves the number, and nobody has measured how far.

It does not mean one clean measurement. The abstract and the conclusion call 97.3 percent a mean cross-validation accuracy. The results table lists the same figure under a test set heading. Those are two different measurements, and the paper never clears it up. So you do not know which one 97.3 is.

It does not mean the model can name the problem. There are three labels. Affected. Healthy. No cactus. The affected label covers both the cochineal insect and fungal rot, and treatments differ. So the app tells a farmer something is wrong. It does not tell them what to do.

It does not mean 97.3 percent on the hard question. One of the three groups is background photos of soil, sky and weeds. All three models got those right 99 percent of the time. An easy group pulls an average up. The hard case is an old scar against fungal rot, because both look rough. That one mix-up caused 68 percent of the small network’s errors.

It does not mean you should ship the most accurate model. MobileViT-XS is 9.3 MB and takes 68 milliseconds per photo. The small network is 4.8 MB and 42 milliseconds, at 89.5 percent. On a cheap phone with little room left, those numbers weigh as much as accuracy. The paper also does not say which device the timings came from.

And it does not mean the app works. The app is built. It speaks Tigrigna and Amharic. It runs offline. The paper reports no session with a real farmer, not one. An accurate model in an app nobody uses is worth nothing.

What it does mean

On photos of the kind this team collected, a 9.3 MB model separates sick pads from healthy ones well. That is a real result, and mixing two designs is why. MobileViT-XS looks at a whole photo at once, so it can see that pest wax comes in clusters while old scars sit alone. It threw out 94 percent of the false alarms the plain networks raised.

The bigger result is not the model. It is the 3,587 photos. Global crop datasets are built from flat leaves such as apple, maize and tomato. A cactus pad is thick, waxy and spined, and models trained on flat leaves do poorly on it. Somebody had to walk out and photograph cactus pads. That is the part nobody else had done.

They also gave the photos away. Code and dataset are public. That is why the 97.3 can be argued about at all.

Why it matters to you

Every AI tool you are offered arrives with one accuracy number on the front. It was measured on examples somebody chose, with labels somebody wrote. It was almost never measured where you plan to use it.

The gap between a paper number and a field number is not a rounding error. It is the whole product. A model at 97 percent in a lab and 70 percent in a village is a different business.

You cannot see that gap in the number. You see it in the questions underneath.

Do this today

Take the last accuracy number anyone showed you. Ask three things. What set was it measured on, and who collected it? How many groups is it sorting into, and is one of them easy? Which two groups does it confuse most, and what does that mistake cost? If nobody can answer the third, the number is decoration.

Just Out Tech explains new research in plain language. This article was drafted with AI assistance and checked by a human against the original source.

What to remember
  • The 97.3 percent came from photos the team shot at university test plots, and the paper reports no session with a real farmer.
  • The same 97.3 percent figure is described as a mean cross-validation accuracy in the abstract and as a test set result in the table, and the paper never resolves the difference.
  • The three labels are affected, healthy and no cactus, so the model can say a pad has a problem but cannot say whether the problem is an insect or a rot.

Questions people ask

Does 97.3 percent mean the app will be right 97 times in 100 for a farmer?

No. It is the score on photos the research team took at test plots at Mekelle University and Adigrat University. A farmer holds the phone differently, in different light, on differently dusty pads. The paper reports no field trial, so nobody has measured how much the number moves.

Can the app tell an insect apart from a fungus?

No. The dataset has three labels only: affected, healthy and no cactus. The affected label covers both cochineal infestation and fungal rot. The model warns you that a pad has a problem, but the paper does not show it naming which problem, and the treatments differ.

Why does the easy background group matter?

Because accuracy is an average across all three groups. All three models were right about background photos of soil, sky and weeds 99 percent of the time. An easy group lifts the average. The hard case is an old scar against fungal rot, which caused 68 percent of the small network's errors.

Can I check the 97.3 percent myself?

Yes. The paper links a public GitHub repository for the source code and says the cactus-fig dataset has been open sourced and is reachable through the project repository and Kaggle. That openness is the strongest thing about this work, because another group can rerun it.

About the author

Mark Alex

Mark Alex is the founder and Managing Director of Real Biz Digital, a technology company operating out of Nairobi since 2018. He works in agentic AI and the Model Context Protocol, AI governance, enterprise software architecture and cybersecurity. He holds an MSc in Mechatronical Engineering from Obuda University in Budapest and a BSc in IT, Forensic Technology and Cybercrime, from USIU-Africa in Nairobi, and has published IEEE conference research on an AI-powered digital twin for greenhouse systems. He is the author of seven books. Between 2020 and 2024 he mentored more than 200 university students and interns in Nairobi. He writes every Just Out Tech article from the original research paper.