Intelligent Vision turns any webcam, IP camera, phone or even your screen into an AI observer. It sees who is there and how they feel, what they hold and what they do, what kind of place it is and what is happening right now — and it keeps learning from every camera in your business.
Every picture from the camera goes through a chain of small, fast neural networks on your PC. Each one adds a layer of understanding; together they turn pixels into people, emotions, objects, behaviours, places and moments.
A pose network finds every person and 17 body joints — head, shoulders, elbows, wrists, hips, knees, ankles. From the joints alone the app knows if someone stands, sits, walks, runs, lies down, raises a hand or falls, even when the face is not visible.
A face detector plus emotion networks read happy, neutral, sad, angry, surprised, fearful; a face-recognition model remembers returning customers, VIPs, staff and blocked persons. The clothes of the day are remembered too, so a known person is recognised at once even with the face turned away.
Two detectors (80 + 601 classes) find cups, bottles, phones, laptops, food, animals, vehicles, weapons, number plates… A crop around each wrist is classified separately: cigarette, lighter, glass, phone, knife. Every doubtful object gets a second opinion from a 1000-class classifier, the shared memory and — when allowed — internet reference photos.
Joints and objects combine into behaviours: drinking (cup to mouth), smoking (repeated hand-to-mouth without a cup), on the phone, eating, reading, working on a laptop, sleeping, talking with someone, fighting, loitering… Each one carries a real confidence: how stable it was, how well the skeleton was seen and whether the trained model agrees.
Every 5 seconds the full frame is classified: cafe, office, classroom, road, sea, gaming screen, kitchen, gym… and every frame gets a moment: "a person beating another", "2 people having breakfast", "nice view: mountains and a river", "driving on a road", "a crowd of 12 people".
Who drank 3 glasses of wine 40 minutes ago, who smoked, who slept — remembered for hours and across visits. Over many visits the app learns a person's temperament: calm, cheerful, moody, agitated, aggressive — or your own word for them.
Rotate through the stack: each layer is a network or a rule set that runs on your PC in a few milliseconds. Nothing leaves the machine except tiny learning features.
The neural networks run on the customer's Windows PC — on its GPU when it has one, on the CPU otherwise. The server never sees video: it receives only compact learning features and small screenshots, keeps the shared memory, trains the models and manages licenses. One small Ubuntu VM serves hundreds of cameras.
GPU build: NVIDIA, AMD or Intel graphics through DirectML (CUDA for NVIDIA), 25+ frames a second with all layers. CPU build: any Windows 10/11 PC, 8-12 fps — enough for a shop or a classroom.
A 2-vCPU VM is enough for hundreds of cameras: it stores features, confirms them by policy, checks doubtful looks against the internet, trains the four models on the CPU in seconds and pushes them out. S3 / MinIO for unlimited pictures.
Open the page on any phone, enter the key, point the camera: frames go to the server, the same AI runs there on the CPU (4 fps per phone), alerts and learning come back. Ideal for a quick check of a site or a temporary event.
Out of the box it knows 600+ things and 40+ behaviours. Everything beyond that it learns from you, from videos, from the internet and from your other cameras — and it trains its own models without a data scientist.
Draw a box on the live picture and type the name — "espresso machine", "our sign", "safety helmet". It is remembered at once, used on screen and in rules, and shared with your other cameras through the central memory.
Browse a movie file, paste a YouTube link or a folder of recordings: the app watches every frame, collects behaviours, faces, objects, scenes and moments, and you label what matters. Great for rare events: fights, falls, robberies.
"Gym floor": based on the cafe mode, behaviours "lifting weights, stretching", things "dumbbell, yoga mat", a sentence to say. Label 10 samples each, press Train — the HUD shows your counters, rules fire on your words.
The memory works through 700 names from the internet before a camera ever sees them, learns what goes with what (a boat in an office needs more proof than a cup in a cafe) and answers questions: "who drank wine this week?", "when was fire last?".
Every visit updates a profile: usual mood, mood swings, fights, smoking, drinking, when they come, how long they stay. The model says "usually cheerful · regular · comes ~18:00" or flags "aggressive · risk 60 %" — and your tags teach it.
Upload a new build in the admin panel; every running exe downloads it in the background, checks its signature and restarts when nobody is in view. Two builds — GPU and CPU — each device fetches its own kind.
Pick a business mode and the app renames its behaviours, loads the fitting rules and watches what matters there — or leave it on Auto: it recognises the place by itself and suggests the mode.
Attentive vs looking down, hand raised, reading / writing, phone in class, sleeping, who is present.
Drinks per person, alcohol counting, "2 people drinking + fast arm movements → possible fight", smoking, sleeping guests, VIPs.
Footfall and dwell time per zone, crowd alerts, loitering, running, unattended bags, lost children (child near no adult), moments across many cameras.
Gate mode: known residents vs strangers, blocked persons, loitering at night, vehicles and number plates, pets, falls of elderly residents, fire and smoke.
Visits and mood per customer, returning guests by name, waiting too long, hands raised for service, table occupancy, "2 people having breakfast", inventory of what is in the room.
Office mode: working, meeting, on the phone, talking, attentive; after-hours presence, blocked persons, weapon detection, fire and smoke, tailgating at doors.
Helmet / no helmet (teach your gear in one click), falls, people in danger zones, vehicles and forklifts near people, smoking near materials, fire.
Number plates read and logged, vehicle counts, people loitering between cars, running, night activity, taxis and vans, distance to the camera on a radar.
Shop mode: browsing vs waiting at the counter, what customers pick up (inventory learns your products), queue length, returning customers, blocked shoplifters, staff presence.
One file, no installer. Double-click it, the control panel opens in your browser, choose the camera, enter your license key. It checks for updates by itself from then on.
Not sure? Take the CPU build — it runs everywhere. Windows SmartScreen may ask on the first start ("More info → Run anyway"); the file is published by JOY SERVICES IOT.
Supported: Windows 10 and Windows 11, 64-bit. The app runs as a normal user, no admin rights needed.
| CPU build (no GPU) | GPU build | |
|---|---|---|
| Processor | Any 64-bit PC; Intel Core i5 / Ryzen 5 (8th gen or newer) recommended for 10+ fps | Any 64-bit PC |
| Graphics | not needed | NVIDIA GTX 1050 / RTX or newer, AMD RX 500+ or Intel Arc / Iris Xe — DirectML, driver up to date. NVIDIA + CUDA 12 for the fastest option |
| Memory | 8 GB RAM (16 GB with many cameras) | 8 GB RAM + 2 GB GPU memory |
| Disk | 2 GB free (models, learning data, snapshots) | |
| Cameras | USB webcam · phone as webcam (Camo, iVCam, EpocCam) · IP / CCTV camera via RTSP · YouTube link · video / movie file · screen share (desktop or one window) | |
| Network | Internet for activation, the shared memory, updates and internet verification; it keeps running for 72 h without it | |
| Speed you can expect | 8-12 fps, all layers on (objects every 3rd frame) | 25-30 fps, everything every frame |
No. The video stays on your PC. What travels to the server is tiny: a 1000-number "look" of an object, a 63-number pose feature, small thumbnails and one full-frame screenshot per minute — the material the shared memory learns from. Phone mode is the exception: there the frames go to the server because the AI runs there.
Free: 1 device, runs 10 minutes then pauses 5 (your data keeps training the shared models). Normal: 3 devices, no pauses, bigger pictures. Pro: 25 devices, largest pictures, priority. Keys are made in the admin panel; the same key can be moved between PCs.
No. The CPU build runs on any Windows 10/11 PC at 8-12 fps, which is plenty for a shop, a classroom or a cafe. A GPU makes it 25-30 fps and lets you run the heavier models on every frame. Either way, no GPU is needed on the server side.
Frontal, well-lit faces of 60 px or more: reliable for happy / neutral / negative; sad vs angry vs fearful is harder for any system. The app says "unclear" instead of guessing on small or turned faces, and you can train a calibration for your camera in a few minutes.
Yes — that is the point. Click a thing and name it; name a scene ("racing game"), a moment ("robbery"), a person ("trouble"); feed it movies. What one camera learns reaches every camera of your business within seconds, and the four models are retrained from it automatically.
USB webcams, phones as webcams, any IP camera or DVR that gives an RTSP stream (rtsp://user:pass@ip:554/…), HTTP / MJPEG streams, YouTube videos and live streams, local video files and your own screen (a game, a movie, a video call).
The exe is published by JOY SERVICES IOT (right-click → Properties → Details) and contains no installer or background service. On the first start choose "More info → Run anyway". Signed builds are on the way, which removes the warning entirely.