AI video analytics tender specification showing CCTV monitoring, detection rate, false alarm rate and acceptance criteria

How to Specify AI Video Analytics in a Tender (So the Vendor Can’t Wriggle Out)

There is a system I think about often. It passed its acceptance test cleanly. Every camera was up, every analytic rule fired during the demo, the signatures went on the page, and the project closed. Six months later the operators had quietly built a habit: click, dismiss, click, dismiss. The alarms still arrived. Nobody read them. On paper the system was working perfectly, and in practice it had been switched off by human beings who simply stopped believing it.

Nothing broke. No component failed. What failed was written months earlier, in a document nobody looked at again — the tender. Getting the AI video analytics tender specification right is not a technical problem that engineers solve after installation; it is a drafting problem that buyers solve before a single camera is ordered. If your specification does not define what “working” means in numbers, under stated conditions, measured on your own site, then the vendor gets to define it for you at acceptance — and they will define it generously.

I’ve spent years on the specification side of public safety procurement. What follows is what I wish more buyers put into their documents. It isn’t vendor-hostile. A good vendor will read these clauses and price them properly. That is exactly the point.

What “AI-powered” actually means in a specification

Short answer: nothing. “AI-powered,” “deep learning,” and “intelligent detection” carry no contractual weight. Only three things are measurable — the detection rate, the false alarm rate, and the conditions under which both were measured. Without the third, the first two are marketing.

Every serious analytics claim collapses into those three components. A vendor who tells you their system is highly accurate has told you nothing until you ask: accurate at detecting what, missing how often, and under which lighting, weather, camera geometry and scene? Accuracy figures travel without their conditions the way a car’s fuel economy travels without the test cycle — technically true, practically useless.

So the first discipline in drafting is negative: strike every adjective from the analytics section and replace it with a number that has a measurement context attached. If a requirement cannot be failed, it is not a requirement.

The two numbers vendors quote separately

Detection performance is always two numbers, not one. There is the probability of detection — of the real events that occurred, how many did the system flag? And there is the false alarm rate — of the alarms the system raised, how many were nothing at all. They pull against each other. Lower the threshold and you catch more real events while drowning in noise; raise it and the noise disappears along with the events you actually cared about.

Vendors know this, and quote whichever number flatters the configuration they’re proposing. A specification that demands one without the other has demanded nothing, because either can be driven to look excellent by sacrificing the other. Both numbers must be required together, in the same clause, measured in the same test, under the same conditions.

The rare event trap

Here is the part that catches experienced buyers, and the reason a “99% accurate” system can be worthless. Perimeter intrusion, abandoned objects, wrong-way movement — these are rare. The maths of rare events is brutally counterintuitive.

Take a modest site. Say the analytics evaluate 10,000 detection opportunities in a week across your cameras — motion sequences, object tracks, people crossing lines. Say two of those are genuine events you care about. Now suppose the system is excellent: it catches both real events, and its false positive rate is just 1%.

QuantityValue
Detection opportunities per week10,000
Genuine events2
False positive rate1% of the ~9,998 non-events
False alarms produced≈ 100 per week
Real alarms produced2 per week
Chance any given alarm is real≈ 2%

A 99% accurate system, on a site with rare events, produces alarms that are wrong roughly forty-nine times out of fifty. The vendor’s number was honest. The operator’s experience is that the system cries wolf all day. This is the base rate problem, and it is not a flaw in the product — it is arithmetic that applies to every detector ever built.

The practical consequence: on a rare-event site, the false positive rate is the specification that matters, and it has to be tight to a degree that sounds unreasonable until you run the multiplication. Run it with your own numbers before you write the threshold. Bring the result to the vendor meeting.

There is a second-order effect worth naming, because it’s what actually kills systems. Research from the security industry has repeatedly found that operators facing a constant stream of false notifications begin ignoring them within a couple of weeks. Once that habit forms, your detection rate is effectively zero regardless of what the analytics are doing. Alarm fatigue is not a training issue you can fix later. It is a design constraint you specify against.

Writing the acceptance test

This is the section missing from almost every tender I’ve read, and the one that does the most work. An acceptance test is not a demonstration. A demonstration is the vendor showing you their system succeeding; an acceptance test is you attempting to make it fail under agreed rules and recording what happened.

  1. Specify a continuous test window, not a demo slot

    Minimum 72 hours of uninterrupted operation with logging enabled. A two-hour demo tells you the software runs. Three unbroken days tell you what the system does at 3 a.m., during shift change, and when the weather turns. State that the window runs continuously and that any interruption restarts it.

  2. Mandate condition diversity

    The window must include night with IR illumination, full daylight, and at least one adverse condition — rain, fog, or reduced visibility. No condition may be waived for convenience. If the test period doesn’t naturally produce the adverse condition, the clause should say the test extends until it does, rather than allowing the requirement to be dropped.

  3. Define the false alarm unit as per camera, per day

    This single line changes the balance of the whole document. A total count across the system lets a vendor hide a catastrophic camera behind forty quiet ones. A per-camera, per-day rate makes every scene accountable on its own terms. Write it explicitly: false alarms shall not exceed [X] per camera per 24-hour period, with no single camera exceeding the threshold.

  4. Set thresholds per scene, not globally

    A loading bay, a fence line facing a public road, and a stairwell are three different statistical problems. One global default cannot be a fair acceptance criterion for all of them. Group cameras into scene classes in the specification and give each class its own numbers.

  5. Test on your footage, not theirs

    The test runs on live streams from the delivered installation on your site. Pre-recorded vendor footage, curated clips, and reference installations are evidence of capability, never evidence of acceptance. Say so in the clause, because otherwise it will be offered.

  6. Write down what happens when the threshold is missed

    Most specifications stop at the number and leave the consequence unwritten, which in practice means the consequence is a meeting. Define it: a remediation period of [X] days, a full re-test under identical conditions, a maximum number of re-test attempts, and what final failure means commercially — withheld payment, extended warranty, partial rejection, whatever your framework allows.

⚠️
Ground truth has to be agreed before the window starts. Who decides an alarm was genuine? If that’s settled after the data comes in, both sides will interpret ambiguous events in their own favour and the test resolves into a negotiation. Name the method and the arbiter in the document.

The clause everyone forgets: degradation over time

Systems pass on day one and fall apart by month six, and the reason is almost never the software. The environment moves. Seasons change the light and the thermal background. Vegetation grows into the frame. Someone installs a new floodlight that throws a moving shadow across a detection zone. A camera mount develops a vibration in wind. A scene gets rearranged — pallets stacked where there was open ground, a fence repositioned, a new access route worn into the grass.

Each of these silently shifts the statistics the system was tuned against. Detection drifts down, false alarms drift up, and both happen slowly enough that nobody can point to the day it broke. If your specification is silent on this, the recalibration burden lands on you by default, and you will be paying for it as a change request.

What belongs in the contract:

ClauseWhat it must state
OwnershipWhose responsibility recalibration is, by name of role, for the full contract term — not “as required.”
Periodic obligationA fixed cadence of performance verification (quarterly is a defensible starting point) measured against the original acceptance thresholds.
Event triggersAny physical change — camera repositioning, new lighting, construction, scene layout change, vegetation management — triggers recalibration and re-verification of the affected cameras.
Standard restoredAfter recalibration, the system must meet the original acceptance numbers, not a renegotiated set. Without this sentence the thresholds erode with every revisit.
ReportingA per-camera performance report at each verification, delivered in writing, retained for the contract term.
“A system that met its specification on the day of handover and nowhere else has not been delivered. It has been demonstrated.”

Camera specification is analytics specification

Software cannot recover a pixel the sensor never captured. This should be obvious, and yet the most common failure I see in tender documents is a beautifully detailed analytics annex bolted onto a camera annex that says almost nothing — resolution, IP rating, warranty, done. The analytics were then evaluated separately from the hardware that feeds them, as if they were independent purchases.

Four hardware parameters do more damage to analytics performance than any model choice:

Pixel density on target. Not sensor resolution — pixels across the object, at the far edge of the detection zone, at the actual mounting geometry. A 4K camera badly placed delivers fewer usable pixels on a person at 40 metres than a well-placed 2MP one. Specify pixels-per-metre at the furthest point of each detection zone and require the vendor to demonstrate it in the design, before installation.

Real night performance versus IR distance on the datasheet. IR range figures describe illumination reaching a surface, not usable detail arriving at the sensor. Beyond a certain distance you get a bright blur, and the analytics see a bright blur. Require night performance to be evidenced at the specified detection distance, on site, as part of the same acceptance window.

Compression. Aggressive bitrate limits smear exactly the fine motion and edge detail that detection depends on, and the damage is invisible to a human reviewing footage on a monitor. Set a minimum bitrate or maximum compression level for cameras running analytics and forbid unilateral reduction for storage reasons after handover.

Frame rate and mount stability. Low frame rates break object tracking; a mount that moves in wind turns a static background into a moving one, which is the single most reliable way to generate false alarms. Specify a minimum analytics frame rate and a mounting rigidity requirement, including pole-mounted cameras under wind load.

ℹ️
A useful drafting habit: for every analytics threshold you write, ask what hardware condition would have to hold for that number to be achievable. If the camera annex doesn’t guarantee that condition, your analytics threshold is unenforceable — the vendor will point at the hardware you specified and be right.

A copy-paste acceptance criteria template

Below is a skeleton you can lift into your own tender and fill in. The bracketed values are yours to set based on your scene classes and your event rate — run the base rate arithmetic above before choosing them. Adjust the language to your own procurement framework; the structure is what matters.

Analytics Performance and Acceptance

Replace every [X] with a project-specific value.
  1. Scene classes. Cameras are grouped into scene classes [list]. All performance criteria below apply per scene class and are evaluated per camera.
  2. Detection rate. The system shall detect not less than [X]% of genuine [event type] events within each scene class, measured over the acceptance window defined in clause 5, against ground truth established under clause 8.
  3. False alarm rate. The system shall generate not more than [X] false alarms per camera per 24-hour period. No individual camera may exceed this figure. System-wide totals or averages shall not be used to demonstrate compliance.
  4. Joint evaluation. Clauses 2 and 3 shall be demonstrated simultaneously, in the same test window, at a single fixed configuration. Threshold or configuration changes between the measurement of detection rate and false alarm rate void the test.
  5. Acceptance window. Continuous operation of not less than [72] hours on the delivered installation, using live streams from the Purchaser’s site. Vendor-supplied recordings and reference sites are not admissible.
  6. Condition matrix. The window shall include, for each scene class: (a) night-time operation with the installed IR/illumination; (b) full daylight; (c) at least one adverse visibility condition — [rain / fog / low light]. Failure to capture any listed condition extends the window until it is captured.
  7. Hardware preconditions. Minimum [X] pixels per metre on target at the furthest point of each detection zone; minimum analytics frame rate [X] fps; minimum stream bitrate [X] Mbps, not to be reduced post-acceptance without written agreement.
  8. Ground truth. Genuine events shall be determined by [method / joint review panel], with the procedure agreed in writing before the window opens. Disputed events are resolved by [named arbiter].
  9. Re-test. If any criterion is not met, the Contractor shall remediate within [X] calendar days and repeat the full window under clauses 5 and 6. A maximum of [X] re-tests is permitted.
  10. Non-acceptance. Failure after the permitted re-tests constitutes non-conformity, with consequences as set out in [contract clause reference].
  11. Ongoing calibration. The Contractor shall verify performance against clauses 2 and 3 every [3] months for the contract term, and shall recalibrate and re-verify affected cameras following any change to camera position, lighting, scene layout, vegetation or surrounding construction. Post-recalibration performance shall meet the original thresholds.
  12. Reporting. Each verification shall produce a written per-camera performance report delivered within [X] days and retained for the contract term.
One caution about published numbers. Almost every “reduced false alarms by 90%” figure in circulation originates from a vendor’s own material, measured on a site you’ll never see, under conditions never stated. Treat those figures as directional at best — an indication that a category of technology can help, never a measurement you can specify against. The only performance numbers that belong in your acceptance clause are the ones produced on your cameras, on your site, during your test window.

What this actually buys you

A document written this way does something subtle at bid stage, before any equipment exists. It changes which vendors want the job. Bidders who intend to install a default configuration and leave will read the per-camera false alarm clause and the recalibration obligation, price the risk honestly, and either walk away or come back with a number that reflects the real work. Bidders who tune scenes properly will price it as ordinary effort, because for them it is.

That sorting happens quietly, at no cost, months before anyone signs. It is the highest-leverage thing a specification does, and it’s available to anyone willing to write six numbers and a consequence clause instead of the word “intelligent.”

Frequently asked questions

The vendor says their system is 99% accurate. Is that a lie?

Probably not. It’s likely a real measurement from a real dataset — just not from your site, your cameras, your lighting or your event rate. As the base rate example above shows, a genuinely 99%-accurate detector on a rare-event site still produces alarms that are mostly wrong. The number isn’t dishonest; it’s irrelevant until it’s reproduced on your installation. Ask what conditions it was measured under, and if the answer isn’t specific, treat it as a claim of capability rather than performance.

Can I just specify a false alarm number and be done?

No — a false alarm threshold alone is trivially satisfiable by desensitising the system until it detects almost nothing. You need it paired with a detection rate, at a single fixed configuration, measured in the same window, with the measurement conditions defined. Any one of those four missing and the clause can be met without the system being useful.

What if the vendor refuses these terms?

That refusal is information, and it’s cheaper to receive at bid stage than at handover. Some pushback is legitimate: a vendor may reasonably decline to guarantee performance on cameras or mounts they didn’t specify, or on scenes you haven’t let them survey. That’s a signal to fix your hardware annex or fund a site survey, not to drop the clause. Refusal to commit to any measurable per-camera figure on an installation they designed themselves is a different signal entirely.

Does any of this change for cloud versus edge analytics?

The performance clauses don’t change at all — detection rate, false alarm rate and conditions are architecture-neutral, and you should resist any suggestion that one deployment model deserves softer numbers. What changes is what you add alongside them. Cloud deployments need clauses on bandwidth dependency, behaviour during connectivity loss, latency from event to alarm, and where footage is processed and stored. Edge deployments need clauses on processing capacity headroom, how models are updated on the device, and whether an update can silently change tuned thresholds — which it can, and which is worth a sentence forbidding it without re-verification.

Written by
Yavuz Yasin Çetinkaya
AI Automation Specialist & Workflow Architect
AI and video surveillance specialist with 16+ years of field experience.

Leave a Reply

Your email address will not be published. Required fields are marked *