Almost every camera system sold in British Columbia is described as having "motion detection". The phrase appears on the quote, in the app, and in the conversation on site. It sounds like a single feature that a system either has or does not have.
It is three different technologies, doing three different amounts of work, and the gap between the first and the third is the difference between a system you trust at three in the morning and a system whose notifications you turned off in March.
Nobody is lying to you. The word is just doing far too much work.
One: pixel-change detection
The oldest and cheapest method compares each frame to the one before it. If enough pixels changed by enough, that is motion.
That is the whole mechanism. It has no concept of an object, a person, a vehicle, or a scene. It knows that the picture is different now, and it treats difference as event.
Consider what makes a picture different at night in a wet climate.
Rain, lit by the camera's own infrared illuminator, turns into bright drifting streaks a few centimetres from the lens. Headlights sweep across a wall. A security light kicks on and the exposure hunts, changing every pixel in the frame at once. A branch moves. A flag moves constantly. Cloud crosses the moon and the whole yard shifts brightness. A spider builds a web across the housing and sits on it, enormous and out of focus, for six hours. Snow. Steam off a roof vent. A cat.
Every one of those is a genuine, correct detection of pixel change. The system is not malfunctioning when it fires on them. It is doing exactly what it was built to do, and what it was built to do is not what you wanted.
Two: VMD with rules and zones
Video motion detection — VMD in most manuals and menus — is the same idea with structure added around it. It is a real improvement and it is worth configuring properly.
The tools are familiar to anyone who has been into a camera's settings. Zones let you mark part of the frame as watched and the rest as ignored, so the road at the top of the image stops generating events. Sensitivity and object-size thresholds require a change to be large enough or strong enough before it counts, which suppresses the spider and much of the rain. Dwell and count rules require the change to persist for a period, or to happen repeatedly, before an alert leaves the system. Line-crossing and intrusion rules add direction and geometry: something must cross this line, this way, or remain inside this shape.
Applied by someone who walks the site and thinks about it, this removes a great deal of noise. A well-tuned VMD setup on a fenced yard is not a bad system.
But look at what is underneath. Every one of those rules is a filter on pixel change. The zone says where the change may happen, the threshold says how much, the dwell says for how long, the line says in which direction. None of them says what.
So the failure mode narrows but does not change character. Headlights sweeping through your zone are large, persistent, directional pixel change. Rain in the illuminator inside the zone is still rain. A tarp working loose on a windy night crosses your line repeatedly, all night, in both directions. Wind in BC does not respect an object-size threshold, and neither does a cardboard box.
And the tuning has a cost that is easy to miss. Every filter you tighten to kill a false alarm also raises the bar for a real one. Push the size threshold up until the raccoon stops triggering, and a person crouched behind a bin may no longer clear it either. You are not removing noise; you are trading sensitivity for quiet, blind to what you gave up.
Three: AI object classification
The third technology asks a different question. Instead of did enough pixels change, it asks what is that.
A classifier is a model trained on a very large number of labelled images to recognize categories: a person, a vehicle, sometimes finer distinctions such as a car versus a truck versus a bicycle, sometimes an animal class. It runs on the video — increasingly on a processor inside the camera itself rather than on a distant server — and for each thing it finds it outputs a category and a confidence score.
The alert rule then sits on top of that output, and this is where the whole benefit comes from. You are no longer saying "tell me when pixels change in this zone". You are saying "tell me when a person is in this zone".
Rain is not a person. Headlights are not a vehicle — the light is not the car, and a classifier looking for a vehicle shape does not find one in a moving patch of brightness. The flag is not a person. The spider on the lens is not a person, no matter how large it looks. The tarp is not a person. All of it is still happening in the video, and none of it produces an alert, because none of it is in a category you asked to be told about.
That is a genuinely different kind of filtering. VMD suppresses noise by describing the noise — its size, its duration, its location. Classification suppresses noise by describing what you want, and everything else falls away without you having to anticipate it. You do not need to have thought of the spider in advance.
This technology is real and it works, and the scepticism in the rest of this article is not scepticism about whether it works.
What alert fatigue actually costs
The reason this distinction matters more than most camera specifications is that the failure it prevents is a human one, and it is terminal.
A system that produces alerts nobody can act on does not stay in a steady state of being annoying. It decays, and it decays in a predictable order.
First, the alerts are read. Someone checks each one, sees rain, and moves on. Then they are skimmed — glanced at, dismissed without opening the clip. Then the notification is muted on one phone, usually at night, usually by the person who most needed it. Then the rule is disabled, or the sensitivity is dropped so far that the system has effectively been switched off while continuing to appear switched on. Then, months later, something actually happens, and the alert either never fired or fired into a channel that no living person still reads.
At that point the equipment is still on the wall, still recording, still on the maintenance schedule, and still on the balance sheet. What it has lost is the only property that made it worth more than a recorder: someone finding out while it is happening.
This is the same argument as the one for thermal on a dark perimeter — a detector is only worth having if people still believe it in month six — arrived at from the other direction. Thermal earns trust by being physically indifferent to the things that fool an ordinary camera. Classification earns it by discarding them after the fact. Both are attacking the same enemy, which is not the intruder. It is the shrug.
What classification does not fix
Here is where a proposal should be read carefully, because "AI" is doing marketing work as well as technical work, and the two are not always separable in a brochure.
It tells you what, not who. A classifier reports that a person is present. It does not report which person. Your night cleaner, your tenant's teenager, a delivery driver at the wrong address, and someone who should not be there all classify identically, because they are all people. Classification removes weather and animals from your alert stream. It does not remove your own staff. If you want alerts that distinguish authorized from unauthorized, that is a scheduling and access-control question — who is expected here, at this hour — and the camera cannot answer it alone.
It does not produce identification. Being told a person is there is not the same as having an image that establishes who they were. That still depends on pixels on the subject, lens choice, lighting, and where the camera is pointed. Axis, describing thermal cameras, makes the underlying distinction in its own product literature: "Unlike conventional cameras, this technology only allows for detection." The detection-versus-identification split is not unique to thermal. Any detector, however clever, tells you that. Producing an image good enough to identify someone is a separate design problem, on a separate camera, with separate requirements.
A confidence figure is not a guarantee. Classifiers output a number expressing how strongly the model matches what it has seen to a category. That number is a property of the model's response to this image, not a probability that the world is a certain way, and it is not a percentage of correctness you can hold anyone to. Confidence relates to conditions — distance, occlusion, unusual pose, heavy rain, an angle unlike the training data, someone lying down or partly behind a vehicle. Push the threshold up and you get fewer false alerts and more missed people. Push it down and the reverse. There is no setting that is simply "accurate", and anyone offering you a single accuracy number for your site, sight unseen, has told you something about their sales process rather than about the equipment.
It is only as good as its categories and its view. A model trained to find people and vehicles will find people and vehicles. Something that matters to you but is not a category — a door left open, a pallet removed, a bicycle taken off a rack — is not detected by a classifier, and asking for "AI detection" does not conjure it. And a classifier still cannot see through a badly placed camera. Too far, too high, too backlit, too low a resolution on target, and the model is guessing on smudges.
The question to ask
You do not need to evaluate a model. You need one question, and it works on any proposal, from anyone:
"When this system alerts me at three in the morning, what has it already ruled out?"
A pixel-change system has ruled out nothing. Something in the picture changed.
A tuned VMD system has ruled out change that was too small, too brief, in the wrong place, or in the wrong direction. Useful, and worth having configured properly, but everything large and persistent inside your zone still reaches you.
A classification system has ruled out everything the model did not place in the categories you asked about — the weather, the animals, the lights, the vegetation, and the things nobody thought to anticipate.
Then ask the follow-up, which is the one that separates a designed system from a delivered box: "What is still in the alert stream that I do not want, and where does that alert go?" An honest answer names something. Your own staff arriving at shift change. Vehicles legitimately using the lane. Deliveries. Whoever is quoting you should be able to say what your remaining noise will be, and what schedule, zone, or arming state deals with it — because an alert nobody reads is the same as no alert at all, whichever of the three technologies produced it.
The short version
"Motion detection" describes three unrelated levels of intelligence. Pixel-change detection reports difference. VMD reports filtered difference. Classification reports objects, and only the objects you asked for.
The third is worth paying for, on most sites, for one reason: it is the only one whose alerts survive contact with a wet winter, and a system whose alerts you still open is the only kind that does anything at all.
Just do not let it be sold to you as more than it is. It tells you something is there and roughly what kind of thing it is. It does not tell you who, it does not hand you an identifying image, and its confidence score is a dial with a cost on both sides — not a promise.
Sources
- Axis Communications, Thermal imaging (technical overview — the detection-versus-identification distinction) — https://www.axis.com/solutions/thermal-imaging
