Computer vision

Make sense of
images and video

Cameras are cheap and everywhere. The hard part is turning what they see into a number your systems can act on, reliably enough that somebody stops checking it by hand.

Rated 4.9 on Clutch across 38 reviews


2-4weeks to a model trained
on your own images
6-10weeks to a deployed
pipeline in production
100%of predictions carry
a confidence score
100mstypical inference at the edge,
on modest hardware

What we build

Vision covers a lot of ground. These are the applications we are asked for most.

Quality inspection

Defects caught on the line at a consistent standard, without depending on how tired the person on shift is.

Detection and counting

Objects, people or vehicles found and counted in a frame, with the count written into your systems rather than a clipboard.

Document capture

Scans and photographs of forms turned into structured fields, including the handwritten ones, with low-confidence cases flagged.

Video analytics

Events found in hours of footage, so a person reviews the ninety seconds that matter instead of the whole recording.

Classification

Images sorted into your own categories, trained on your own examples rather than a generic label set.

Visual search

Finding the matching item from a photograph, across a catalogue that changes faster than anyone can tag it.

How it works

A vision model is only as good as the images it was trained on. Most of the effort goes into the dataset, not the architecture.

Your side

What you capture

  • Existing cameras or scanners
  • Archived images and footage
  • Photographs from the field
  • Your labels and categories
What we build

The vision pipeline

  • Annotation and dataset build
  • Model training and validation
  • Confidence thresholds
  • Human review for the uncertain
  • Drift monitoring
Behind it

Where it runs

  • Edge device on site
  • Your cloud or ours
  • Batch over an archive
  • API into your systems

How we work

Vision projects fail on data, not on models. We front-load that.

1

Look at the images

A week with your actual footage, working out whether the signal is even present and what the lighting, angles and edge cases really look like.

2

Build the dataset

Annotation against a labelling standard your team agrees with. This is the slowest step and the one that decides the outcome.

3

Train and validate

A model trained on your data and measured on a held-out set, reported as precision and recall rather than a single accuracy number.

4

Deploy

To the edge device or the cloud, whichever the latency and connectivity actually require, with confidence thresholds set with you.

5

Monitor for drift

Cameras move, seasons change and products get redesigned. We watch for the accuracy decay and retrain before it becomes a problem.

Built by a team that ships

The AI layer is new. The engineering underneath it is not. These are products we designed, built and still maintain.

Planable Omniconvert Wolfpack Digital OPEN social CANGO Mobility Zerotak TaskManager Bookster Life in Codes HCT Envision Tickbird
VerityPanel Market research platform

Data collection at scale

A digital market study and data-collection platform built for fast, reliable research at scale.

Product development · Frontend
DIY design space Kitchen design

Design it before you buy it

DIY Design Space is an interactive tool that lets customers design and visualise their dream kitchen online.

UI/UX · Frontend
Mobility Mobility

Fleets, tracked

CANGO Mobility builds fleet and mobility software for operators managing vehicles at scale.

Product · Web development

The choices that matter

Vision has more deployment constraints than most AI work. These are the ones that shape the build.

Fine-tuned open models, vision language models, or classical CV

Classical methods still win on well-controlled industrial images, and cost far less to run. We check that before reaching for a network.

Your team, our team, or model-assisted

Model-assisted labelling cuts the effort substantially once a first model exists, but somebody who knows the domain has to set the standard.

On the device, on a server in the building, or in the cloud

Decided by latency, connectivity and whether the footage is allowed to leave the site. Often the answer is on-device.

How precision and recall trade off against each other

A single accuracy figure hides which mistake you are making. We set the threshold with you, knowing which error is more expensive.

The cameras you have, new capture, or industrial sensors

We start with what you already have. Better lighting frequently beats a better model, and costs less.


“Inventiff. is easy to work with, and nothing is a problem for them.”

Sean Williams · CTO, HCT Concierge Verified review on Clutch

Got an idea? Let’s make it real.

Tell us the short version

This could be the first step towards a new and successful collaboration. A one-line idea and a finished spec are both fine — tell us the problem, the deadline you’re working to and what’s in your way.

We reply within one working day.

Prefer another way to talk?

Frequently asked questions

Usually not. We start with the footage and hardware you already have. Better lighting or a small change of angle frequently beats a better model, and costs a fraction as much.

It depends on how varied the task is, but far fewer than teams expect for a narrow, well-controlled problem. We give you a realistic figure after looking at your actual images, not before.

Yes. Where latency or connectivity requires it, the model runs on a device at the site, and only the results are sent onward. That is often also the answer when the footage is not allowed to leave the premises.

We report precision and recall separately rather than a single accuracy figure, because it hides which mistake you are making. The threshold is set with you, knowing which error costs more in your operation.

Cameras move, seasons change and products get redesigned. We monitor for the decay and retrain before it becomes a problem, which is why a vision system is an ongoing arrangement rather than a delivery.

Two to four weeks to a model trained on your own images, and six to ten weeks to a deployed pipeline, with most of that time going into the dataset rather than the model.