SE-VISION · Release 2026.1

A camera estate that finally tells you something.

Most organisations already have hundreds of cameras and almost no information. SE-VISION is a full video management system — devices, recording, failover, evidence — with a GPU analytics engine built into the core rather than sold as a licence pack on top.

0Per-frame inference
0Streams per GPU node
0Analytics rule types
0Raw video over the WAN
3 rules armed · 0 false alarms in 24 h

The problem

Analytics projects rarely fail on model accuracy.

They fail on the six weeks after go-live, when the alarm queue fills with wind-blown rubbish and reflections, an operator mutes the notifications, and the system quietly becomes a very expensive recorder again.

01

Alarms nobody trusts

A tripwire that fires forty times a night trains your operators to ignore it. By month two the rule is disabled, and the incident it was bought for goes unseen.

02

Analytics sold by the licence

Detection is priced per camera, per rule, per year. Enabling a behaviour on a new camera becomes a purchase decision, so coverage stays where the budget landed in year one.

03

Bandwidth as the ceiling

Central analytics needs every stream at the core. Remote sites on a modest link either get nothing or get a degraded stream the model was never trained on.

What it does

A VMS you would keep, and analytics you would actually leave switched on.

The device and recording layer has to be genuinely enterprise-grade before the clever part matters. Both halves ship together.

01 · Devices & recording

Boring, dependable plumbing

ONVIF Profile S, G and T, RTSP and RTP transport, H.264, H.265 and MJPEG. Dual-stream by default: the high-quality stream is recorded, a low-resolution stream feeds analytics, so neither job compromises the other. When a network gap ends, edge storage on the camera is pulled back and stitched into the timeline automatically.

  • Auto-discovery and bulk provisioning by subnet
  • Edge storage retrieval after a link failure
  • PTZ, presets, tours, audio and dry-contact I/O
  • Recording server pools with failover recording
02 · The analytics engine

Inference at the edge, events at the centre

Detection, tracking and re-identification run on a GPU node next to the cameras. Only events cross the network — a class, a confidence, a track identifier, a bounding box and a thumbnail — typically a few kilobytes where a stream would be megabits. A remote site on a modest link gets the same analytics as headquarters.

  • Under 60 ms per frame, 128+ streams per GPU node
  • Person, vehicle, ANPR/LPR, PPE, abandoned object
  • Horizontal scale-out by adding nodes, not licences
  • Models updated without touching the recording layer

Why this matters commercially: analytics is part of the platform, not a per-camera add-on. Turning a behaviour on for another fifty cameras is a configuration change, not a quotation.

Diagram: a frame passes through decode, inference, tracking, rules and event publication, with per-stage latency adding to under 60 milliseconds; only events leave the edge node.
Frame pipeline · latency budget per stage
03 · Rules

Tuning is the product

Regions, tripwires, direction, dwell, occupancy and queue length are drawn on the scene by an operator, not written by an integrator. Every rule carries suppression controls — minimum persistence, size gates, schedule windows, and masks for the tree that moves in the wind — because the difference between a useful system and a muted one is entirely in these settings.

  • Draw-on-video rule authoring with live preview
  • Persistence, size and schedule gates per rule
  • Replay a rule against last week's recording before arming it
  • Per-rule alarm statistics so you can see what is noisy
04 · Forensic search

Find the person, not the timestamp

Search a time window by attribute — a person in a dark jacket, a white van, a bag left alone — and get candidate clips across every camera. Cross-camera re-identification stitches a subject's path together so an investigator follows one track instead of scrubbing eleven timelines.

  • Attribute and appearance search across cameras
  • Stitched multi-camera tracks with per-hop confidence
  • Operator confirms or rejects each hop; both are logged
  • Export with watermark, hash and chain-of-custody record
Diagram: a subject tracked on camera three is matched on camera seven at 0.93 confidence and camera eleven at 0.87, forming one stitched track.
Cross-camera re-identification · appearance, not identity
05 · Operator experience

Built for a control room at 03:00

Alarm queue with acknowledge, escalate and disposition codes. Map-based navigation so an operator finds the camera by where it is, not by what somebody named it in 2019. Video wall layouts driven from the same console, and a mobile client for the supervisor who is walking the floor rather than sitting at it.

  • Alarm queue with disposition and handover notes
  • Floor-plan navigation with camera field-of-view cones
  • Multi-monitor and video wall layout control
  • Mobile client with live view and clip sharing
06 · Privacy & governance

Designed for a regulator to read

Retention is set per camera, not globally, so a car park and a staff area can differ. Redaction blurs faces and plates on export by default and requires a named privilege to disable. Every export is hashed and logged with who requested it, why, and what was released.

  • Per-camera retention with automatic expiry
  • Default-on redaction for export, privileged override
  • Immutable access log covering live view and playback
  • Privacy masks burned in at the edge, before recording

What we do not sell: we do not ship watchlist facial recognition as a turnkey feature. Face detection for counting and redaction, yes; matching a face against an identity database is a decision with legal and ethical weight that belongs to you and your counsel, not in a default configuration.

Architecture

Edge does the thinking, the core does the remembering.

Recording stays close to the cameras. Analytics runs beside it. Only events and requested clips travel — which is what makes multi-site deployment affordable.

SE-VISION runtime architecture Cameras connect to a site node running recording and GPU inference; events flow to the management server, while operator, wall and mobile clients connect to the core. SITE · EDGE CORE CLIENTS IP cameras ONVIF · RTSP · H.265 GPU inference DETECT · TRACK · RE-ID 128 STREAMS / NODE Recording server LOCAL STORAGE POOL RETENTION PER CAMERA Rule evaluator ROI · DWELL · LINE Management server DEVICES · USERS · RULES CONFIG DISTRIBUTION Event store SEARCHABLE · INDEXED THUMBNAILS ONLY Evidence service EXPORT · HASH · REDACT CHAIN OF CUSTODY Operator client ALARMS · PLAYBACK Video wall LAYOUT CONTROL Mobile client LIVE · CLIP SHARE Northbound PSIM · MQTT · WEBHOOK THE WAN CARRIES EVENTS AND REQUESTED CLIPS — NOT CONTINUOUS VIDEO A SITE KEEPS RECORDING AND ALARMING IF THE LINK TO THE CORE IS DOWN
Node

GPU inference

One or more per site. Add a node to add streams; no per-camera analytics licence.

Node

Recording server

Pooled, with failover recording so a host loss does not create a gap.

Service

Event store

Indexed detections and tracks. Holds thumbnails, never the underlying video.

Service

Evidence service

The only path that can export video, and the only one that can lift redaction.

What the operator sees

The same frame, before and after the engine.

Drag the handle. The left side is what a recorder gives you; the right is what the analytics layer adds on top of the identical stream.

An illustration, not a capture

This is a rendered mock-up of the overlay, not a screenshot of a real deployment — we will not present synthetic footage as evidence of accuracy. What it shows faithfully is the information model: a class, a confidence, a persistent track identifier and a rule state, drawn over the stream the operator is already watching.

  • Boxes carry a stable track ID across frames, not per-frame detections
  • Confidence is shown so an operator can judge a marginal call
  • Rule state (loitering, 12 s) is separate from detection
  • Overlays are a client-side layer — the recording stays unmodified

On accuracy claims: we will not quote a single accuracy figure for your site until we have run against your footage. Anyone who quotes one before that is quoting a benchmark, not your car park at dusk.

Specifications

The page your security integrator reads.

Architecture & components
Core services
Management server, recording server pool, GPU inference node, rule evaluator, event store, evidence service
Runtime
Containerised; Kubernetes or Docker Compose; Ubuntu LTS or RHEL hosts
Acceleration
NVIDIA GPUs with CUDA; T4, A2, A10 and L4 class validated; CPU-only fallback for low channel counts
Storage
Local block storage or S3-compatible object storage for archive tiering
Clients
Windows and Linux desktop, web operator client, iOS and Android mobile
Devices & protocols
Standards
ONVIF Profile S, G and T; RTSP, RTP, RTCP; ONVIF events
Codecs
H.264, H.265/HEVC, MJPEG; up to 4K per channel
Streams
Dual-stream per camera: recording stream and analytics stream, independently configured
Camera features
PTZ with presets and tours, two-way audio, dry-contact input and relay output, edge storage retrieval
Provisioning
Subnet discovery, bulk credential apply, per-model profile templates
Analytics catalogue
Detection
Person, vehicle with class, two-wheeler, bag, face (detection only), PPE items
Recognition
ANPR/LPR with regional plate formats; cross-camera person re-identification by appearance
Behaviour rules
Region entry and exit, tripwire with direction, loitering and dwell, speed, crowd density, queue length, occupancy counting, abandoned and removed object, safety-zone breach
Suppression
Minimum persistence, object size floor and ceiling, schedule windows, static masks, per-rule statistics
Rule authoring
Drawn on live video by an operator; replay against recorded footage before arming
Capacity & performance
Inference latency
Under 60 ms per frame; 50 ms typical at 1080p and 12 fps analytics rate on a T4-class GPU
Streams per node
128+ analytics streams per GPU node, depending on rate and model mix
Recording channels
Not licence-capped; bounded by disk throughput and network
Event bandwidth
Kilobytes per event; a site with 200 cameras typically stays well under 1 Mbit/s upstream
Scale-out
Add GPU nodes for analytics, recording servers for channels; both horizontal
Evidence, privacy & compliance
Export
Watermarked, hashed, with a signed manifest; player bundled for offline review
Chain of custody
Every export logged with requester, justification, scope and hash; log is append-only
Redaction
Face and plate blurring on by default for export; disabling requires a distinct privilege
Retention
Per camera and per zone with automatic expiry; legal hold overrides expiry
Privacy masks
Applied at the edge before recording, so masked regions are never written to disk
Data protection
Deployment patterns for India's DPDP Act and for GDPR; DPIA support material provided
Deployment, licensing & support
Deployment models
On-premise single site, multi-site with edge nodes, or hybrid with a cloud management plane
Air-gapped
Fully supported; model updates delivered as signed offline bundles
Licensing
Per recorded channel plus per GPU node for analytics; perpetual or subscription
Support
Standard included; Enhanced and Critical tiers — see support tiers
Tuning service
Camera survey and rule tuning available — see vision deployment services

Interoperability

Systems we already speak to.

Protocols and system classes we have built adapters for. Naming them states interoperability, not partnership.

ONVIFS · G · T RTSP / RTPTransport H.264 / H.265Codecs Access controlDoor events Intrusion panelsZone state ANPR barriersPlate gating PSIMNorthbound MQTTEvent bus WebhooksCustom hooks SAML / OIDCIdentity NVIDIA CUDAAcceleration S3-compatibleArchive tier PrometheusMetrics SNMPNOC alerts

Deployment

Three ways to run it.

Single site, on-premise

One rack: recording, GPU inference and management together. Air-gapped if required.

  • Simplest to operate and audit
  • No outbound dependency
  • Typical for one campus or plant

Multi-site with edge nodes

A node per site keeps recording and analytics local; the core sees events only.

  • Works over modest WAN links
  • Sites keep alarming if the link drops
  • The usual choice for retail and logistics estates

Hybrid management plane

Video never leaves your sites; configuration, users and event search live in your cloud.

  • Central administration across regions
  • Video remains on-premise by design
  • Data residency configurable per region
0Per-frame inferenceMeasured at the inference node
0Streams per GPU node1080p at 12 fps analytics rate
0Typical tuning periodGo-live to stable alarm rate
0Upstream for 200 camerasEvents only, no continuous video

Before you ask

The questions that decide the deal.

What accuracy will we get?

We do not know yet, and neither does anyone else quoting you a number. Accuracy depends on your camera angles, lens choice, lighting at the worst hour of the day, and what you consider a true positive. Give us a week of footage from your five hardest cameras and we will report measured precision and recall on that footage, including the cases we get wrong. That number is worth something; a datasheet figure is not.

Do we have to replace our existing VMS?

No. We regularly run the analytics engine against an incumbent VMS's RTSP streams, so you keep your recording estate and add the intelligence layer. If you later want to consolidate, migration runs camera group by camera group with both systems recording in parallel. Both paths are supported and priced differently — we will tell you which is cheaper for your estate.

How does this compare to Milestone XProtect?

On device support, recording architecture and evidence handling we are aiming at the same class of system, and XProtect has a far longer track record and a deeper partner ecosystem. Where we differ is that analytics is in the core rather than a per-camera licence pack, rules are tuned by your operators rather than by an integrator, and you get direct access to the engineers who wrote it. If you need a global vendor's support footprint across forty countries, that is a real reason to buy XProtect.

What about facial recognition and privacy law?

We ship face detection — used for counting and for automatic redaction — but not watchlist matching against an identity database as a turnkey feature. Where a lawful basis genuinely exists, that becomes a scoped engagement with your data protection officer involved, not a checkbox in a configuration screen. For India's DPDP Act and for GDPR we provide deployment patterns, retention defaults and DPIA support material.

What happens in week one versus week six?

Week one is noisy — expect a false-alarm rate that would be unacceptable long term. That is the tuning period, and it is the work: persistence gates, size floors, masks and schedules against your real scenes. By week six a well-tuned perimeter rule should be producing single-digit nightly alarms. Anyone who tells you it is accurate on day one has not tuned it against your site.

What does SE-VISION deliberately not do?

It is not an access control system and not an intrusion panel. It consumes their events and correlates against video; it does not unlock doors or arm zones. It also will not make a poorly positioned camera work — if a lens is looking into the sun at 17:00, the honest fix is a bracket and a survey, and we would rather quote that than sell you a model that cannot help.