Computer vision in business is, alongside chatbots, the AI technology with the widest gap between what gets demoed at a trade fair and what survives a night shift on a factory floor. In the demo, the camera spots the faulty part in two tenths of a second. On the real line, an operator moves a lamp, sunlight comes through the loading bay door, and the same model starts flagging good parts as defects. The technology didn't fail: the project did, because it mistook an algorithm for an installation.
This guide is about the part the demo leaves out. What computer vision actually is and how it differs from "putting up an AI camera", which cases are mature today in manufacturing, logistics and retail, what you need before buying hardware, how much a realistic project costs and takes, what GDPR forces you to consider when the camera points at people, and when the honest answer is that it isn't worth it.
What is computer vision, and how is it different from an AI camera?
Computer vision is the set of techniques that let a system pull useful information out of images or video: counting objects, classifying them, locating them, measuring them, reading text or detecting that something is off. The input is pixels; the output is a structured piece of data another system can consume.
That last sentence is the whole point, and it separates a project that pays for itself from an anecdote. A camera that "detects" but triggers nothing — no alarm, no rejection on the line, no entry in the ERP — is an expensive monitor. The value isn't in the recognition; it's in what happens after recognising.
It helps to distinguish four types of task, because their difficulty and cost differ enormously:
- Classification. Is this image a good part or a faulty one? The simplest, and usually the first to work.
- Detection and localisation. How many boxes are there and where is each one? Labelling costs more, because you have to mark positions, not just categories.
- Segmentation and measurement. How much surface does the stain cover, how long is this weld? Demands camera precision and calibration, and it's where most projects stall.
- Reading and extraction. Licence plates, codes, labels, delivery notes. The most commercially solved of the four, and often no custom model is needed at all.
A word on vocabulary: when a supplier says "state-of-the-art AI" about a box counter, trust the model more than the proposal. Around 80 % of industrial vision problems are solved with architectures that have been around for years. The difference is made by lighting, angle, lens cleanliness and who maintains the system in September — not by the novelty of the algorithm.
A computer vision project isn't won in the model. It's won in the lighting, the labelling and the process that consumes the alert. Anyone selling you just the model is selling you a third of the project.
Which computer vision cases work in business today?
The ones that meet three conditions at once: the object being observed is repetitive, capture conditions can be controlled, and there's a clear decision waiting on the other side. This table covers the ones we're asked about most and what each demands.
| Case | What it solves | Critical condition | Decision it triggers | |---|---|---|---| | In-line quality control | Detecting surface, assembly or labelling defects | Fixed lighting and real defect samples | Automatic rejection or operator alert | | Counting and inventory | Counting units, pallets or SKUs in a warehouse | Stable angle and separable objects | Stock adjustment in the ERP | | Plant safety | Detecting restricted-zone access or missing PPE | Clear legal basis and privacy policy | Alert to the shift supervisor | | Traceability and reading | Reading codes, plates, batches or delivery notes | Print quality and focus | Automatic inbound/outbound record | | Store analytics | Footfall, queue times, cold zones | Anonymisation at the camera itself | Staffing changes, layout redesign |
Three caveats that prevent surprises.
Quality control. The problem is never recognising the good part: it's that you have so few bad ones. A line with 0.5 % defects produces two examples per four hundred photos, and a model needs considerably more. That's why serious projects start by collecting and classifying defects for weeks, or by flipping the problem around: learn what normal looks like and flag deviations from it.
Counting and inventory. The return is quick when it replaces manual counts, and poor when it competes with a barcode scanner that already works. Before looking at cameras, check whether the problem is capture or simply that nobody records the movement. Often the right answer isn't vision but automating the process that generates the record.
Safety and people analytics. Technically among the most mature; legally the most delicate. More on that below.
What do you need before installing the first camera?
Five things, and only one of them is technological.
- A concrete decision and its owner. "We want to see what happens on the floor" isn't a case. "We want to automatically reject parts with burrs before packaging, and the line supervisor owns it" is.
- Controllable capture conditions. Constant lighting, fixed camera position, stable background and the object always at a similar distance. It's boring and it's 40 % of the success. A project under variable daylight costs twice as much and delivers half.
- Real images, not catalogue ones. You need photos of your product, your line and your dirt. And you need failure examples. If they don't exist yet, the project's first deliverable is a capture and labelling setup, not a model.
- A written acceptance criterion. What share of defects must be caught and how many false alarms are tolerated. Without that number, the final discussion becomes an opinion, and opinion always loses to "Pete used to do this and he never missed one".
- Whoever consumes the output. PLC, ERP, WMS, shift dashboard or an email. If the detection doesn't land in a system already in use, it'll be abandoned within three weeks. It's the same pattern behind why AI pilots fail.
One note on the asymmetric cost of errors, which almost nobody raises up front. In quality control, letting a defect through doesn't cost the same as rejecting a good part. In safety, a false alarm burns the shift's trust in two days. That balance can be tuned, but it has to be decided with the business in the room, not left at a library's default threshold.
How much does a computer vision project cost and how long does it take?
A first, well-scoped case at a mid-sized company fits in three or four months. This is the sequence we follow and where the effort actually goes.
| Phase | What happens | Output | |---|---|---| | Weeks 1-2 | Define the case, acceptance criterion and current baseline | One-page brief with an owner and numbers | | Weeks 3-6 | Build the capture rig: camera, optics, lighting, position | Real, reproducible images of the process | | Weeks 5-9 | Labelling and first model; validation on fresh data | Does it meet the criterion? Answered with figures | | Weeks 10-13 | Integration with PLC, ERP or dashboard; shift training | Detection inside the actual workflow | | Weeks 14-16 | Supervised operation and threshold tuning | Decision: scale, adjust or stop |
Three uncomfortable budget realities. First: hardware is usually the smallest line item. A decent industrial camera, its optics and lighting run into hundreds or a few thousand euros per station; integration and tuning cost several times that. Second: labelling always gets paid for, in money or in your people's hours, and it deserves a named owner and a calendar slot. Third: budget for maintenance. Change the packaging supplier, change the sheen of the plastic, and the model starts to miss. Without someone reviewing metrics and retraining, the system degrades quietly.
Always compare against the baseline: what share of defects escapes human inspection today and what does each escape cost. If nobody has measured that, measuring it is the project's first job. It's the same discipline we apply in predictive analytics: with no baseline, any result looks good.
What does GDPR require when the camera points at people?
The moment a camera captures an identifiable person, you're processing personal data — even if all you want is a headcount. That triggers obligations best resolved before installation, not after the first complaint.
- Legal basis and notice. Legitimate interest is the usual route for workplace safety, and it requires a documented balancing test. Plus visible signage and information to staff and their legal representatives.
- Genuine minimisation. If you need to count people, don't store faces. Serious systems process on the device itself and emit only the number. That's anonymisation applied to video, and it cuts risk more than any contract clause.
- Retention periods. Video surveillance footage has short statutory retention limits as a general rule. Keeping it "just in case" is the most common breach and the easiest to avoid.
- Impact assessment. Systematic monitoring of publicly accessible areas or of workers usually requires one. Doing it isn't bureaucracy: it's what forces you to write down what you do with the footage.
- Bounded workplace use. Detecting a missing helmet is safety control; measuring individual work pace crosses into employee monitoring, with its own rules and far more conflict.
On top of that sits the European AI framework, which classifies by risk and treats biometric identification with particular severity. If your case only counts objects, the impact is limited; if it identifies people, check which category applies before signing. We summarise it in our guide to the EU AI Act for SMEs, and the practical rule is simple: the less you identify, the fewer obligations you accumulate and the fewer arguments you'll have with the works council.
When is computer vision NOT worth it?
Ruling a case out early preserves the budget for the one that works. Clear signals:
- The volume doesn't justify it. If fifty parts are inspected a day, a person is cheaper, more flexible and never needs retraining.
- The product changes constantly. Very short, very varied runs mean relabelling forever. Vision pays off where there's repetition.
- A better signal already exists. A weight sensor, a barcode scanner or a scale solve many "vision problems" at a fraction of the cost and with no ambiguity.
- The scene can't be controlled. Outdoors, shifting light, overlapping objects and a camera somebody moves every week: the system will spend more time misaligned than working.
- Nobody will act on the alert. If the detection lands on a dashboard nobody watches, don't build it. Build it once it's clear who acts and how fast.
In quite a few of these cases the right next step isn't more technology but tidying up the record of what already happens. The same conclusion as always: data before AI.
Frequently asked questions
Do I need a lot of images to train a computer vision model?
Fewer than people assume for classification, and considerably more for rare defects. A few hundred examples per category with good capture conditions goes a long way; the bottleneck is usually gathering enough failure examples, not good-product ones.
Does computer vision replace quality inspectors?
In practice it reassigns them. The system does the repetitive sweep across 100 % of parts and flags the doubtful ones; the person resolves edge cases and keeps the criterion consistent. Projects framed as full replacement tend to underdeliver and generate internal resistance.
Can I use the CCTV cameras I already have?
Sometimes, for counting and footfall. For quality control almost never: resolution, optics, position and lighting are designed to watch, not to measure. Reusing unsuitable hardware is the saving that ends up costing most.
Does the video have to go to the cloud?
Not always, and often it shouldn't. Processing on the device cuts latency, network cost and risk surface — something you'll appreciate when you get to the cybersecurity review. The cloud makes sense for training and for consolidating metrics, not necessarily for inference.
If you're weighing up computer vision and aren't sure whether your case is mature — or whether the real bottleneck is the lighting, the labelling or the process meant to consume the alert — that's exactly what the audit resolves: looking at your line, your images and your systems to come out with a prioritised case, a measurable acceptance criterion and the discards in writing. If you'd rather sanity-check it in half an hour first, let's talk.
Shall we apply it to your case?
The 360° AI Audit turns these ideas into a concrete plan for your company: three weeks, fixed price and the full picture of your AI before spending a euro.
See the 360° Audit→ Let's talk↗