Catching Shelf Gaps in Kenyan Retail with AI Image Recognition

A field rep finishes a call at a duka off a busy road in Nairobi, holds up a phone and photographs the shelf behind the counter. It takes four seconds. By the end of the day there are sixty more like it; by the end of the month, several thousand. Nobody has opened the folder.

That is the ordinary state of shelf photography in Kenyan field sales. Capture is cheap; reading is not. A supervisor who reviewed every photograph would spend more time on pictures than on the team. So the images sit there as proof that a visit happened, while the questions they could have answered — was the fast-mover in stock, did a competitor take the eye-level slot — go unanswered.

AI image recognition is the obvious response: let a model read the photograph and produce the answers. It works, and it is not magic. It works well in parts of Kenyan retail, partially in others, and in a good number of outlets not at all. This post is about telling those apart before you commit a budget.

A shelf photo with abstract AI bounding boxes highlighting gaps and misplacements

What Image Recognition Actually Does to a Shelf Photo

The process has two steps. The model draws a box around each thing it believes is a product, then matches each box against a trained catalogue of pack images — this SKU, that variant, that pack size. It returns a list: what it found, where it sat in the frame, and how confident it is.

Everything after that is arithmetic on the list. Share of shelf is your boxes divided by all boxes. A gap is an expected SKU that produced no box. Planogram compliance is the list compared against a layout you defined. Misplaced facings are boxes in the wrong region. No downstream number is more reliable than the detection beneath it.

Two consequences follow. A model only recognises packs it has been trained on, so somebody must maintain an image catalogue for the packs actually sold in Kenya — including the sachets, refill packs and small formats that fill duka shelves and rarely appear in a global pack library. And an unrecognised pack is not an absent pack. A redesigned carton comes back as "unknown"; if your process reads "unknown" as "not stocked", the team spends the month chasing gaps that were never there.

The Four Questions a Shelf Photo Can Answer

Almost every shelf-image programme answers four questions, and they are not equally difficult.

  • Availability. Is the pack present at all? The easiest question and usually the most useful — an out-of-stock on a fast-mover costs a sale today and teaches the shopper to accept a substitute tomorrow.
  • Share of shelf. How much visible space is yours against the rest of the category? Dependable where product is faced in rows, unreliable where it is stacked.
  • Planogram compliance. Does the shelf match the specified layout — sequence, block, shelf height? Meaningful only where a planogram exists and the fixture can hold it.
  • Misplaced facings. Is your pack in the wrong block, or someone else's in a slot you paid for? Useful in organised aisles, close to meaningless in a stacked duka.

A fifth question is one people expect a photograph to answer and it usually cannot: price. Shelf-edge labels are common in supermarkets and mini-marts, rare in dukas and kiosks, where price is spoken rather than displayed. Price stays a manual field on the visit form.

Why a Duka Shelf Is Not a Planogram

The planogram idea assumes a fixture: bays of known width, shelves at known heights, a defined range, space to face product in rows. Most of Kenyan retail has no such fixture, and a programme designed as though it does produces numbers nobody in the field recognises.

A typical duka is a small room stocked to the ceiling. Cartons are stacked rather than faced, so one visible pack may front two dozen more. Sachets hang in strips from the door frame. Fast-movers sit behind the counter within the shopkeeper's reach; slower lines are in the back or under the counter. A kiosk may serve through a hatch, with no shelf visible from the customer side at all. An agrovet has a different layout again — heavy containers low down, treatment lines behind glass, a store room doing the real work.

In that setting "planogram compliance" is not a hard metric, it is a category error. The useful questions are simpler: is the must-stock list present, are the strips still hanging, is the pack somewhere a customer can see it, is it facing forward at all. Reframe around availability and visibility and the programme describes something real.

Light, Angle and the Limits of the Camera

Even where a shelf can be photographed, the physics work against you. Many dukas are deep, narrow and lit by one bulb, the brightest thing in the frame being the doorway behind the rep — a backlit image with the shelf in shadow. Plastic-wrapped multipacks throw reflections. There is rarely room to stand back far enough for a whole section, so the rep shoots one bay at a time, or at an angle that distorts the sizes the model is comparing. Older handsets soften detail, and upload compression removes exactly what separates one SKU from its sibling: a variant colour band, small print on a sachet.

The fixes are unglamorous. Specify the shot — how many photographs, from which position, covering which section. Run a quality check on the device before submission, so a dark or blurred image is rejected while the rep is still in the shop rather than flagged three days later. Then protect consistency: the same outlet shot the same way month after month yields a trustworthy trend even when absolute counts are imperfect.

A Morning on a Nakuru Beat

Picture an ordinary Tuesday on a Nakuru route. The rep starts before eight at a wholesaler near the town centre, then works outward along a residential road: a run of dukas, a kiosk or two, a mini-mart, and later an agrovet.

At the first duka the shelf sits behind the counter and the shopkeeper hands goods over rather than letting anyone through. The rep photographs what is visible from the customer side. Three-quarters of the relevant packs are in frame; the rest are behind the shopkeeper's shoulder, so the system will report a gap on a line sitting just out of view. The second stop is a kiosk serving through a hatch, where there is no shelf photograph to take at all.

The mini-mart three hundred metres on is a different world: open aisles, faced product, shelf-edge labels, daylight through a glass front. The photograph is clean, the model reads it well, and share of shelf and misplaced facings both mean something. So does the agrovet in the afternoon, once the rep learns to shoot the two bays carrying the range rather than the whole room.

Out of twenty-odd calls, perhaps a third produce an image a model can read confidently, a third something partial, and the rest nothing usable. That is not a software failure, it is the channel. A programme that budgets for it works; one that promised head office full shelf coverage does not.

Measuring Share of Shelf When Nothing Is Faced

Share of shelf is a facings metric. It assumes products stand in rows, each visible unit representing roughly equal space. Where a duka stacks cartons the assumption breaks both ways: one visible pack may front a full case, while a competitor with three loose units at the counter looks dominant on almost no stock.

Report what the photograph genuinely supports instead:

  • how many of your distinct variants are visible at all;
  • whether any hold a prime position — counter front, eye level, beside the till;
  • whether hanging strips or a shelf talker are still in place;
  • how those answers have moved since the last visit.

Movement is the honest signal. An absolute share-of-shelf figure from a stacked duka invites arguments nobody can settle. The same outlet, the same shot, four visits running, showing your second variant disappearing — that a territory manager can act on.

Where Supermarket Aisles Change the Maths

None of this is an argument against the technology. In supermarkets, hypermarkets, mini-marts and petrol-station forecourts every assumption behind image recognition holds: standard fixtures, faced product, lighting designed for shopping, aisles wide enough to capture a full bay, a real planogram to compare against. There, gap detection, share of shelf, block sequence and misplaced facings are all measurable, and the saving is real: a merchandiser can photograph a category far faster than they can count it, and counting is where errors creep in.

Kenya's weighting means this is the smaller share of volume. Informal outlets carry most of the trade; modern trade is the secondary motion. That argues for scoping the programme to the formats where it works, staffing those visits accordingly, and running a lighter availability check everywhere else — not for one uniform process across a channel that is not uniform. Photography inside a retailer's store is also at the retailer's discretion, so agree it before the team starts shooting.

Connectivity, Power and the Cost of Photographs

Photographs are heavy. Five to ten images per call across a full beat is a real data load, repeated daily by every rep, on a network that is dependable in the towns and patchy further out. Power is a second constraint: Kenya has seen supply interruptions, including a nationwide outage in 2026, and electricity costs remain high. Neither a charged phone nor a well-lit shop is a safe assumption on an upcountry beat.

So capture offline and queue locally, compress on the device, and sync when the network returns or the phone reaches Wi-Fi at the end of the day. Show a sync status per visit so nothing sits silently unsent. And decide the shots per call and the resolution deliberately: data bundles are paid in KSh by somebody — the company, or worse, the rep out of pocket — and that decision is what your monthly bill and your battery life look like.

Where These Programmes Go Wrong

Most disappointments here are not model failures. They are design failures, and they repeat.

  • Buying recognition before defining the norm. A gap exists only relative to a must-stock list for that outlet class. Without a clean outlet master, the system tells you what is on the shelf but not what is missing from it.
  • Treating detections as facts. Every detection carries a confidence value. Low-confidence and unknown results need handling as "look again", not silent conversion into out-of-stocks.
  • Letting the pack catalogue rot. Packs are redesigned, sizes added, sleeves appear. If nobody owns the image catalogue, quality decays month by month — the field notices long before head office, and trust goes first.
  • Using the score as a stick. Tie a rep's rating to a shelf score and you will get shelf scores: restacking before the photograph, shooting the best bay, shooting the same shelf twice. That is a predictable response to an incentive, not a discipline failure.
  • Reporting activity instead of outcome. "Photos captured" is not a result. Gaps detected, gaps closed and days to close are.
  • Closing the loop too late. A gap flagged after the rep has left is worth a fraction of one flagged while they are still in the outlet. If detection cannot run in the moment, it must become a task on the next visit.

And the largest: expecting the photograph to replace the conversation. An image will not tell you that the shopkeeper is holding back orders over an unresolved credit note, or that the outlet dropped a line when the price crossed what customers on that road will pay. Those are the reasons behind most gaps, and only a person who asks gets them.

The Human Check That Still Matters

The most valuable thing a rep adds to a detected gap is a reason, recorded in the moment. Is the shelf empty because the delivery did not arrive, because the shopkeeper is short of cash until the weekend, because credit is at the limit, because the line does not sell here, or because stock sits unopened in the back? Those five gaps look identical in a photograph and call for five different responses.

A short reason code at the point of capture turns a detection into something a manager can act on. It also builds a record of how often the system's reading matched what the rep found — which is how you learn your own error rate rather than accepting one produced somewhere else.

Supervisors should sample too. A regular manual review of a random set of images against the system output — a handful per territory each month — keeps everyone honest about how the model performs on Kenyan outlets, on the current pack range, on the phones the team actually carries.

How 1Channel Supports Shelf Image Work

1Channel's retail execution and merchandising tools cover the capture and workflow side — the part that must be dependable before any analysis is worth running.

  • Photo capture inside the store visit, with GPS coordinates and a timestamp on each image, tying a submission to an outlet, a rep and a moment.
  • Multiple shots per checkpoint with a reference image alongside, so the rep sees the intended shelf beside the real one.
  • On-device quality checks, so a dark or blurred photograph is rejected before the rep leaves.
  • Offline-first capture with a local queue and a per-visit sync status, for beats where the network drops.
  • Bypass capture with a recorded reason when a duka is shuttered, an owner refuses a photograph, or a store does not permit photography.
  • Outlet master and per-class stocking norms, so "missing" is defined against what that outlet is meant to carry.
  • Multi-level audit review, scoring and exception routing, with roll-ups by route and territory and drill-down to the image.

Where outlet formats support it, shelf images can be analysed automatically for availability and shelf-space questions; where they do not, the same visit workflow carries a manual availability check, so one process covers the whole route. Which outlets belong in which group is a judgement that comes from walking the beat.

Key Takeaways

Image recognition on shelf photographs is a real tool with a defined range. Scope it to where it works, stay candid about the rest, and it earns its place on a Kenyan beat.

  • Availability beats compliance. In dukas and kiosks, measure presence and visibility of the must-stock list rather than a planogram score the fixture could never satisfy.
  • Scope by store format, not by territory. Supermarkets, mini-marts, forecourts and agrovets support a full read; much of the informal channel supports a partial one or none, and the plan should say so upfront.
  • Detections are estimates, not facts. An unrecognised pack is not an absent one, so route low-confidence results to a human rather than into out-of-stocks.
  • Capture discipline decides everything downstream. A specified shot, an on-device quality check, offline queueing and consistency over time matter more than the analysis applied afterwards.
  • The reason code is worth more than the gap. Delivery, cash, credit, demand and stock in the back room look identical in a photograph and need entirely different responses.

Start with the outlets where a clean photograph is possible, define the must-stock list before you measure against it, and give the field a fast way to say what the picture cannot. Handled that way, image recognition takes a slow counting job off the team and leaves them the part of the visit only a person can do.

Insights

Want to get more insights? Click on a topic below