AI camera analytics · runs on your PC

Cameras that understand people, things and moments.

Intelligent Vision turns any webcam, IP camera, phone or even your screen into an AI observer. It sees who is there and how they feel, what they hold and what they do, what kind of place it is and what is happening right now — and it keeps learning from every camera in your business.

Windows 10 / 11 64-bitNo server GPU — the AI runs on your PCFree plan to start
live · 3 people · scene: cafe
happy 92%
wine glass
drinking · 2 min
Ravi · visit 5 · usually cheerful
now: 2 people having dinner
600+things it recognises out of the box, and it learns yours
40+behaviours: sitting, drinking, smoking, phone, sleeping, fighting, fallen…
4models that train themselves from all your cameras
0server GPUs needed — every PC does its own thinking
How it works

One frame, six looks — thirty times a second.

Every picture from the camera goes through a chain of small, fast neural networks on your PC. Each one adds a layer of understanding; together they turn pixels into people, emotions, objects, behaviours, places and moments.

Camera webcam · RTSP · phone YouTube · screen share 1 · People & pose17 joints per person 2 · Face & emotionwho · happy / angry / sad 3 · Objects in viewand what is in the hands 4 · Behaviourdrinking · phone · fight 5 · Scene & momentcafe · road · "lunch for 2" 6 · Memory, alerts and your panel who did what and when · a person's temperament over visits · alerts: fight, fall, fire, smoking, VIP, blocked person voice · your own rules · phone / web panel · API and webhooks to your systems + compact learning features go to the central memory, so every camera of your business gets smarter together
1

People and their skeletons

A pose network finds every person and 17 body joints — head, shoulders, elbows, wrists, hips, knees, ankles. From the joints alone the app knows if someone stands, sits, walks, runs, lies down, raises a hand or falls, even when the face is not visible.

2

Faces, emotions, identity

A face detector plus emotion networks read happy, neutral, sad, angry, surprised, fearful; a face-recognition model remembers returning customers, VIPs, staff and blocked persons. The clothes of the day are remembered too, so a known person is recognised at once even with the face turned away.

3

Objects, and what is in the hands

Two detectors (80 + 601 classes) find cups, bottles, phones, laptops, food, animals, vehicles, weapons, number plates… A crop around each wrist is classified separately: cigarette, lighter, glass, phone, knife. Every doubtful object gets a second opinion from a 1000-class classifier, the shared memory and — when allowed — internet reference photos.

4

Behaviours from body + objects

Joints and objects combine into behaviours: drinking (cup to mouth), smoking (repeated hand-to-mouth without a cup), on the phone, eating, reading, working on a laptop, sleeping, talking with someone, fighting, loitering… Each one carries a real confidence: how stable it was, how well the skeleton was seen and whether the trained model agrees.

5

The whole picture: scenes and moments

Every 5 seconds the full frame is classified: cafe, office, classroom, road, sea, gaming screen, kitchen, gym… and every frame gets a moment: "a person beating another", "2 people having breakfast", "nice view: mountains and a river", "driving on a road", "a crowd of 12 people".

6

Memory that behaves like memory

Who drank 3 glasses of wine 40 minutes ago, who smoked, who slept — remembered for hours and across visits. Over many visits the app learns a person's temperament: calm, cheerful, moody, agitated, aggressive — or your own word for them.

How it detects an image

Six layers of understanding on one frame.

Rotate through the stack: each layer is a network or a rule set that runs on your PC in a few milliseconds. Nothing leaves the machine except tiny learning features.

pixels · 1280×720 frame
pose · 17 joints per person
faces · emotion · identity
objects · hands · text
behaviours · rules · memory
scene · moment · alerts
  • Frame — captured from any source at up to 30 fps; the AI thread processes the newest frame while the overlay stays smooth.
  • Pose — YOLO-pose finds people and joints in ~20 ms on a GPU, ~120 ms on a CPU; fake people (posters, mannequins, reflections) are filtered by joint plausibility and by "never moves".
  • Faces — YuNet + emotion networks with quality gating (too small, blurred or turned faces say "unclear" instead of guessing); SFace embeddings for recognition; per-camera calibration you can train.
  • Objects — COCO + Open Images detectors alternate; verification by a second classifier, the memory of your cameras, and internet photos; OCR reads number plates, phone numbers and signs.
  • Behaviours — geometric rules on the skeleton + nearby objects, refined by a trained behaviour model; business modes (cafe, office, classroom, shop, society…) rename them for your place and add rules like "2 people drinking alcohol → say it and tell the manager".
  • Scene & moment — full-frame classification maps 1000 classes to places; people's behaviours plus the scene and the time of day become moments; a scene model and a moment model trained on your data take over once they are confident.
INTELLIGENT VISION · 06:14 FPS 27 AI 24/s SCENE cafe #1 Ravi · happy 91% · 12 min drinking (wine glass) · 2x earlier wine glass 88% #2 Visitor · on phone · 4 min hand to mouth 3x → smoking? What the app decided on this frame — and why Person #1 is Raviface embedding matched (0.71) · visit 5 · usually cheerful· came at 18:02, stays ~40 min …and is drinkingwrist near the mouth + "wine glass" in the hand crop, verifiedby the second classifier and your cameras' memory → alcohol Person #2 is on the phone, maybe smokingphone-sized object at the ear · repeated hand-to-mouthwithout a cup → "smoking?" (a trained classifier settles it) Scene: cafe · moment: "2 people at a table, drinks"whole-frame classifier (espresso maker, dining table…)+ objects in view + the scenes you named Confidence is real, not decorationlabel stable 3 s (50 %) + skeleton quality (20 %) + model agrees (20 %)→ 0.5 for a glimpse, 0.8 for a sure rule, 0.98 when the model confirms
Why no server GPU

Every PC does its own thinking. The server only remembers.

The neural networks run on the customer's Windows PC — on its GPU when it has one, on the CPU otherwise. The server never sees video: it receives only compact learning features and small screenshots, keeps the shared memory, trains the models and manages licenses. One small Ubuntu VM serves hundreds of cameras.

Cafe · Windows PCGPU build · 27 fpsall 6 layers run herevideo never leaves the PCpanel: localhost:8765 Shop · older PCCPU build · 8-12 fpssame models, CPU onlyRTSP camera in the shopfree plan: 10 min on / 5 off Phone / browsercamera only, no installAI runs on the server CPU4 frames / s per phone Production server · Ubuntu VM no GPU · SQLite / Postgres · S3 or disk ◆ central memory of the business: looks,scenes, moments, behaviours, profiles ◆ auto-integration: confirm · internetcheck · curiosity · consolidate ◆ 4 models: behaviour · objects ·scenes · moments — trained, pushed ◆ licenses, plans, run-time metering ◆ releases: every exe self-updates Admin panel + websitedevices · licenses · learningmemory · training · releases All other cameraswhat one camera learns,every camera knows withinseconds (10 s heartbeat) features + thumbnails (KB, not video)looks · models · scenarios · updates Data that travels: a 1000-number "look" of an object + a 96-640 px thumbnail, a 63-number pose feature, a full-frame screenshot per minute, events. Never the video stream.

Runs on the customer's PC

GPU build: NVIDIA, AMD or Intel graphics through DirectML (CUDA for NVIDIA), 25+ frames a second with all layers. CPU build: any Windows 10/11 PC, 8-12 fps — enough for a shop or a classroom.

The server remembers and trains

A 2-vCPU VM is enough for hundreds of cameras: it stores features, confirms them by policy, checks doubtful looks against the internet, trains the four models on the CPU in seconds and pushes them out. S3 / MinIO for unlimited pictures.

Phone mode when there is no PC

Open the page on any phone, enter the key, point the camera: frames go to the server, the same AI runs there on the CPU (4 fps per phone), alerts and learning come back. Ideal for a quick check of a site or a temporary event.

How it learns and trains

It gets smarter every day — mostly by itself.

Out of the box it knows 600+ things and 40+ behaviours. Everything beyond that it learns from you, from videos, from the internet and from your other cameras — and it trains its own models without a data scientist.

1 · You teachclick a thing → name itname a scene / a momenttag a person: "crazy", "vip" 2 · Videos & photosmovies, YouTube, your recordingsevery frame is learnedphoto folders / zips per label 3 · It watchesverified looks · stable behavioursnew words it can namescenes, moments, profiles 4 · Central memoryconfirm by policy · internet checkcuriosity: learns names it never sawconsolidate like sleep 5 · Trainbehaviour modelobject · scene · momentCPU, seconds 6 · pushed to every camera in seconds — the next frame already uses it Accuracy is shown honestly: cross-validated, per class, next to the rule-based baseline. You decide the ratios in the admin panel (auto-merge confidence, internet match, how often seen, cycle hours). Free plan: the app runs 10 min, pauses 5 — and everything it saw trains the shared models. Normal / Pro: more devices, bigger pictures, no pauses.

Teach by clicking

Draw a box on the live picture and type the name — "espresso machine", "our sign", "safety helmet". It is remembered at once, used on screen and in rules, and shared with your other cameras through the central memory.

Learn from movies and videos

Browse a movie file, paste a YouTube link or a folder of recordings: the app watches every frame, collects behaviours, faces, objects, scenes and moments, and you label what matters. Great for rare events: fights, falls, robberies.

Your own scenarios

"Gym floor": based on the cafe mode, behaviours "lifting weights, stretching", things "dumbbell, yoga mat", a sentence to say. Label 10 samples each, press Train — the HUD shows your counters, rules fire on your words.

Curiosity and context

The memory works through 700 names from the internet before a camera ever sees them, learns what goes with what (a boat in an office needs more proof than a cup in a cafe) and answers questions: "who drank wine this week?", "when was fire last?".

Person memory

Every visit updates a profile: usual mood, mood swings, fights, smoking, drinking, when they come, how long they stay. The model says "usually cheerful · regular · comes ~18:00" or flags "aggressive · risk 60 %" — and your tags teach it.

Self-updating software

Upload a new build in the admin panel; every running exe downloads it in the background, checks its signature and restarts when nobody is in view. Two builds — GPU and CPU — each device fetches its own kind.

Use cases

One app, every kind of place.

Pick a business mode and the app renames its behaviours, loads the fitting rules and watches what matters there — or leave it on Auto: it recognises the place by itself and suggests the mode.

Classroom / school

Attentive vs looking down, hand raised, reading / writing, phone in class, sleeping, who is present.

attention %hand raisedphonefight

Bar / pub

Drinks per person, alcohol counting, "2 people drinking + fast arm movements → possible fight", smoking, sleeping guests, VIPs.

drinks 3xalcoholfightVIP arrived
crowd: 14

Shopping mall

Footfall and dwell time per zone, crowd alerts, loitering, running, unattended bags, lost children (child near no adult), moments across many cameras.

crowdloiteringdwell timerunning

Society / residential

Gate mode: known residents vs strangers, blocked persons, loitering at night, vehicles and number plates, pets, falls of elderly residents, fire and smoke.

strangerplate ABC-1234fallfire
espresso

Cafe / restaurant

Visits and mood per customer, returning guests by name, waiting too long, hands raised for service, table occupancy, "2 people having breakfast", inventory of what is in the room.

visit 5 · happywaiting 8 minhand raisedmeal
working 46 min · meeting 3

Building / office

Office mode: working, meeting, on the phone, talking, attentive; after-hours presence, blocked persons, weapon detection, fire and smoke, tailgating at doors.

workingmeetingweaponafter hours

Construction site

Helmet / no helmet (teach your gear in one click), falls, people in danger zones, vehicles and forklifts near people, smoking near materials, fire.

helmetfalldanger zoneforklift
MH 12 AB 1234loitering

Parking

Number plates read and logged, vehicle counts, people loitering between cars, running, night activity, taxis and vans, distance to the camera on a radar.

plate readvehicles 12loiteringradar
browsing 4 min · picked: bottle

Shops

Shop mode: browsing vs waiting at the counter, what customers pick up (inventory learns your products), queue length, returning customers, blocked shoplifters, staff presence.

browsingqueue 4picked: bottleblocked person
Download

Latest build. Pick the one for your PC.

One file, no installer. Double-click it, the control panel opens in your browser, choose the camera, enter your license key. It checks for updates by itself from then on.

FASTEST

Windows · GPU build

For PCs with a graphics card: NVIDIA, AMD or Intel (DirectML). 25+ fps with every layer on.
⬇ Download IntelligentVision-gpu.exe
version –size –released –
sha256 –

Windows · CPU build (no GPU)

For laptops and older PCs without a graphics card. Same features, 8-12 fps.
⬇ Download IntelligentVision-cpu.exe
version –size –released –
sha256 –

Not sure? Take the CPU build — it runs everywhere. Windows SmartScreen may ask on the first start ("More info → Run anyway"); the file is published by JOY SERVICES IOT.

1Downloadthe build for your PC — one .exe, nothing to install.
2Run itthe first start unpacks the models (10-15 s); the control panel opens in your browser.
3Pick the camerawebcam, phone (Camo / iVCam), RTSP, YouTube, a video file or your screen.
4ActivateCloud & license → server URL + key. Free, Normal or Pro plan.
5Choose the placeor leave Auto: it recognises the scene and suggests the mode. Teach it your things.

System requirements

Supported: Windows 10 and Windows 11, 64-bit. The app runs as a normal user, no admin rights needed.

CPU build (no GPU)GPU build
ProcessorAny 64-bit PC; Intel Core i5 / Ryzen 5 (8th gen or newer) recommended for 10+ fpsAny 64-bit PC
Graphicsnot neededNVIDIA GTX 1050 / RTX or newer, AMD RX 500+ or Intel Arc / Iris Xe — DirectML, driver up to date. NVIDIA + CUDA 12 for the fastest option
Memory8 GB RAM (16 GB with many cameras)8 GB RAM + 2 GB GPU memory
Disk2 GB free (models, learning data, snapshots)
CamerasUSB webcam · phone as webcam (Camo, iVCam, EpocCam) · IP / CCTV camera via RTSP · YouTube link · video / movie file · screen share (desktop or one window)
NetworkInternet for activation, the shared memory, updates and internet verification; it keeps running for 72 h without it
Speed you can expect8-12 fps, all layers on (objects every 3rd frame)25-30 fps, everything every frame
Questions

Straight answers.

Does my video go to a server?

No. The video stays on your PC. What travels to the server is tiny: a 1000-number "look" of an object, a 63-number pose feature, small thumbnails and one full-frame screenshot per minute — the material the shared memory learns from. Phone mode is the exception: there the frames go to the server because the AI runs there.

What do the plans mean?

Free: 1 device, runs 10 minutes then pauses 5 (your data keeps training the shared models). Normal: 3 devices, no pauses, bigger pictures. Pro: 25 devices, largest pictures, priority. Keys are made in the admin panel; the same key can be moved between PCs.

Do I need a GPU?

No. The CPU build runs on any Windows 10/11 PC at 8-12 fps, which is plenty for a shop, a classroom or a cafe. A GPU makes it 25-30 fps and lets you run the heavier models on every frame. Either way, no GPU is needed on the server side.

How accurate is the emotion detection?

Frontal, well-lit faces of 60 px or more: reliable for happy / neutral / negative; sad vs angry vs fearful is harder for any system. The app says "unclear" instead of guessing on small or turned faces, and you can train a calibration for your camera in a few minutes.

Can it learn things that are not in any model?

Yes — that is the point. Click a thing and name it; name a scene ("racing game"), a moment ("robbery"), a person ("trouble"); feed it movies. What one camera learns reaches every camera of your business within seconds, and the four models are retrained from it automatically.

Which cameras and streams work?

USB webcams, phones as webcams, any IP camera or DVR that gives an RTSP stream (rtsp://user:pass@ip:554/…), HTTP / MJPEG streams, YouTube videos and live streams, local video files and your own screen (a game, a movie, a video call).

Windows says the file is unknown / a scanner flags it.

The exe is published by JOY SERVICES IOT (right-click → Properties → Details) and contains no installer or background service. On the first start choose "More info → Run anyway". Signed builds are on the way, which removes the warning entirely.