Building a Reliable Dog-Image Check

Building a Reliable Dog-Image Check

The problem

In the app, an owner reporting a missing dog uploads a photo, and the app turns it into a lost-dog poster.

If the photo is a cat, a bottle, or a screenshot, the poster fails at its first job.

So we wanted one small check before creating it:

“This doesn't look like a dog. Choose another photo, or use it anyway.”

It sounds like a one-line feature.

It isn't.

There are several ways to answer “is this a dog?”, and the differences only become obvious once you look at the mistakes, latency, unusual images, and what happens when the question changes from “is this a dog?” to “what breed is this?”

We compared three approaches on the same images, using the same decision rule.

Three ways to answer one question

The first was EfficientNet-B0, a small image classifier running directly on the device.

It has roughly 5 million parameters and was trained on ImageNet-1k, where 118 of the 1,000 classes are dog breeds.

The model doesn't actually give us a single “dog” probability. It gives us one probability for every class.

So we add the probabilities of all 118 dog breeds.

If it says:

Pug        30%
Bulldog    30%
Everything else 40%

we don't really care that it couldn't decide between pug and bulldog.

For our feature, the useful answer is simply:

60% dog.

The second approach was Amazon Rekognition.

We send it the image and get labels with confidence scores and bounding boxes. If it returns Dog 98%, that becomes our dog score. If there is no Dog label, the score is zero.

The third was Llama 4 Scout 17B through Amazon Bedrock.

Instead of giving it a fixed set of classes, we asked it a question and constrained the response to JSON:

{
  "has_dog": true,
  "dog_count": 1,
  "confidence": "high"
}

We converted that answer into the same 0–1 score used by the other two approaches and ran it at temperature 0.

Three very different systems.

The test had to be harder than the demo

A benchmark full of obvious dog photos doesn't tell you much.

So we built a dataset of 284 images: 142 dogs and 142 non-dogs.

Our own photos deliberately included difficult cases — a lion, a bear cub, a kitten, a dog far away on a lake, and dogs wearing costumes.

AFHQ added close-up faces of dogs, cats, wolves, foxes, lions and tigers. Oxford-IIIT Pet gave us full-scene photos of dogs and cats.

We left out Stanford Dogs and ImageNet because the models had already been trained on them. Testing on the same distribution would make the numbers look better than they really were.

We then split the images into a tuning set of 140 and a test set of 144.

The tuning set was used only to choose the threshold. The test set was kept aside for the final measurement.

Near-duplicate photos were kept in the same split too, so two almost-identical photos couldn't end up on opposite sides and quietly leak information between tuning and testing.

The threshold had one important rule

We didn't just pick 0.5 and call it a day.

For each approach, we tried the possible cut-offs on the tuning set and chose the threshold that rejected the most non-dogs while rejecting no more than 5% of real dogs.

That second part mattered more.

In a missing-dog flow, accepting a bad photo is annoying.

Rejecting someone's actual dog because “the AI doesn't think this is a dog” is a much worse failure.

The cloud services also received exactly the same image bytes. Each image was EXIF-rotated, resized to at most 1024 pixels and encoded once as JPEG. The benchmark ran from inside AWS so the measured latency was closer to the path a production backend would actually see.

We cached every API response as well, so a failed benchmark run didn't mean paying for the same requests again.

The first result was almost a tie

On the final test set, we had 73 dogs and 71 non-dogs.

EfficientNet-B0 reached 97.2% accuracy.

Rekognition reached 95.8%.

Llama 4 Scout reached 97.9%, with none of the 73 real dogs rejected.

The latency told a different story. Rekognition had a median of 167 ms, while the LLM was around 999 ms. The on-device model had no network round trip at all.

The accuracy difference looked impressive at first, but the test set was small. With only 73 dogs, the confidence intervals for the false-rejection rates overlap.

So for the basic dog vs not-dog question, all three approaches were viable.

The interesting part wasn't the leaderboard.

It was the mistakes.

The mistakes had patterns

Every one of the LLM's six mistakes was a wolf or a fox.

Rekognition had the same general problem. It also accepted several wolves and foxes as dogs.

That wasn't random. It tells us something about what these systems consider visually similar.

Rekognition had another interesting behaviour. It could return a strong label such as Cat 100% for a kitten and still attach a weaker Dog label to the same image. A bear cub could get similar treatment.

So simply checking:

if "Dog" exists:
    accept

isn't necessarily enough.

You need to look at the competing labels too.

It also missed some real dogs when something else dominated the image. A Yorkshire Terrier being held by its owner came back primarily as a person. A beagle wearing glasses and a towel was identified as a hound, but didn't receive a Dog label.

The on-device model had a different advantage. It did better on wild animals because ImageNet contains separate classes for wolves and foxes.

But even it had odd failures. A hairless cat was classified as Mexican hairless, which happens to be a hairless dog breed. A Shiba Inu was classified as dingo, which isn't among the 118 domestic-dog classes we were summing.

Every error was explainable once we looked at what the model had actually been trained to know.

The shortcut that looked reasonable

We also tried an obvious shortcut.

We already had a 120-breed dog classifier, a ViT fine-tuned on Stanford Dogs.

Why not use it as the dog detector?

The idea was simple: if it isn't confident about any breed, call it not a dog.

It reached 84.7% accuracy, but rejected 25% of real dogs.

Lowering the threshold didn't fix the problem. It just traded false rejections for false acceptances. A lion became a chow. A bear cub became a Pomeranian.

The problem wasn't the threshold.

It was the label space.

Every class the model knows is a dog breed. Even when you give it a bottle, it still has to distribute its probability across those breeds.

There is no “not a dog” class.

A mixed-breed dog creates the opposite problem: it can look like several breeds at once, so none of the individual breed scores may become high enough.

The general lesson is simple:

A classifier can only reason within the label set it was trained on.

Then the question changed

Once the basic check worked, the obvious next question was:

If it is a dog, can we tell the owner what breed it is?

This time we compared the two cloud approaches using a dataset we collected ourselves: 69 dog photos across 18 breeds.

We removed non-dogs, screenshots, mixed-breed labels, duplicates and photos containing multiple dogs. Four photos whose labels looked doubtful were set aside.

For Rekognition, we used the highest-confidence breed-level label under Dog, while excluding broad groups such as Hound and Terrier and non-breed labels such as Dog Bed.

For the LLM, we asked for the breed, a confidence level and two alternatives.

Synonyms counted as the same breed — Yorkie and Yorkshire Terrier, for example.

This time, the gap was much bigger

Rekognition reached 48% top-1 accuracy.

The LLM reached 88%.

Top-3 accuracy was 49% for Rekognition and 94% for the LLM.

Both systems were perfect on globally common breeds such as Labrador, Golden Retriever and Beagle.

The difference appeared in the less well-covered breeds.

Rekognition correctly identified 0 of 23 Shih Tzus in our test. The LLM identified all 23.

Rekognition also had no correct answer for the Indian native dogs in the set.

The LLM's mistakes were mostly harder pairs: Pug versus French Bulldog, Indie versus Border Collie, German Shepherd versus Husky. In half of its incorrect predictions, the correct breed was already one of its two alternatives.

The latency gap remained, though. Rekognition had a median of 187 ms. The LLM was at 1,141 ms.

The confidence was useful too

The LLM gave us one more thing: a confidence level.

When it said high, it was correct 56 out of 58 times — about 97%.

When it said medium, it was correct only 5 out of 11 times.

That is actually useful at the product level.

We don't have to pretend the model knows something when it doesn't.

If confidence is high, we can show the predicted breed.

If it isn't, we can ask the owner to confirm it.

The model's uncertainty becomes part of the UI instead of something we hide from the user.

The bill for each approach

The on-device model buys you independence. The image never leaves the phone, it works without a network connection, and there is no API latency. But that independence comes with an app-size cost, a mobile inference pipeline, model compression, preprocessing requirements and device-specific performance. Its vocabulary is fixed too; changing what it understands means shipping another model.

Rekognition buys you predictable infrastructure. There is nothing to train or host, latency is low, and the bounding boxes give us information beyond a simple yes or no. But its taxonomy becomes our problem. Foxes and wolves can receive Dog, secondary labels can appear on cats and bears, and breed coverage is uneven. The raw labels need careful post-processing before they become a product decision.

The multimodal LLM buys you flexibility. Counting dogs, ignoring toys or drawings, or asking a different visual question can be handled by changing the prompt rather than retraining a model. It was also the strongest approach in our tests. But the price is latency, model availability, token quotas, defensive parsing and prompt-dependent behaviour. Every meaningful prompt change needs to be evaluated again.

The validation should fail open

There is one decision we'd make regardless of the model.

If the validation service errors or times out, accept the photo.

This check is a quality guard, not the core emergency flow.

We don't want a missing-dog report to fail because an AI service happened to be unavailable.

The model should help us catch bad inputs.

It shouldn't become the thing standing between the owner and the report.

What the experiment actually taught us

Test the model where it is uncomfortable. Easy dog photos don't tell you much. A lion, a bear cub, a kitten or a distant dog can reveal more about a model than hundreds of obvious examples.

Keep tuning separate from evaluation. We used one set of images to choose the threshold and another to measure the final result. Otherwise, you're effectively testing the model against an answer key you helped create.

Look at the raw output. Our first breed run under-scored Rekognition because Dog Bed at 93.7% was being selected above Beagle at 47%. Fixing the parser moved the result from 46% to 48%. The conclusion didn't change, but we only knew that because we looked at what the model actually returned.

Know the label space. A breed classifier can't properly express “not a dog.” A general image classifier can't suddenly recognize a breed it was never trained to recognize. The model's vocabulary puts a boundary around the questions it can answer.

Measure latency where the code actually runs. A cloud API measured from a laptop isn't necessarily the latency a production backend will see. An on-device model that feels fast on a developer phone may behave very differently on a lower-end device.

Make experiments resumable. Caching the API responses meant a crash didn't mean another paid run from the beginning. A small engineering decision, but one that makes repeated experiments much less painful.

Limitations

The test sets were modest: 144 images for detection and 69 for breed identification. Small differences shouldn't be treated as statistically significant.

The breed labels were assigned while collecting the dataset rather than verified by an expert, and Shih Tzu made up about a third of the breed set.

The full-size EfficientNet model was evaluated on a server CPU. A compressed mobile version still needs to be tested on actual phones, especially lower-end devices.

And the LLM results are prompt-dependent. Change the prompt and the behaviour can change too.

Conclusion

"Is this a dog?" looked like one small product requirement.

It turned out to be three separate engineering decisions:

Where does the model run?

What does the model actually know?

How do we turn its answer into a product decision?

For the basic dog-vs-not-dog check, all three approaches performed well enough in our test. The bigger differences appeared when we asked the models to do more, especially breed identification.

The useful conclusion isn't that one model is universally better.

It's that the benchmark score is only one part of the decision.

Test the model on your own data.

Look at what it gets wrong.

Measure the latency where your code actually runs.

And understand what the model was trained to know before asking it a question outside that boundary.

The problem should choose the model, not the other way around.

Read more

Preventing Unsafe Reconfiguration While Subsystems Are Active

Preventing Unsafe Reconfiguration While Subsystems Are Active

Introduction Modern embedded products are becoming increasingly modular. A single hardware platform may support: * Multiple sensor configurations * Different communication modules * Optional accessories * Replaceable compute modules * Expandable peripheral boards This flexibility improves product scalability, but it introduces a hidden reliability challenge: What happens when someone tries to change the system configuration

By Vinayak M K