This website uses cookies

Read our Privacy policy and Terms of use for more information.

London-based GVTLabs is emerging from stealth with a consumer platform to make video searchable. But instead of searching frame by frame, as existing products do, its AI-driven platform can search by actions, movement and physical context. It will launch first as a consumer product, but its underlying technology has implications for the development of Physical AI. 

Its founder, Nick Halstead, is literally one of the ‘original gangsters’ of the Web 2.0 era. He demo’d his first startup to me on a sofa at the launch of TechCrunch Europe in 2007. It was an app that worked with Twitter called TweetMeme, the first site containing what became the ‘retweet’ feature, before it was sold to Twitter in 2010. After exiting a couple of other startups since then, he’s back with what he thinks could be a significant new play in the world where AI meets video. 

AskGVT, is designed to search across entire video libraries and return the exact moment that answers a question, including questions about what someone is physically doing on screen.

“You could ask it, and it finds the exact video clip for something,” Halstead told me. “It can also understand someone gesticulating.”

“If you ask, ‘show me how to use this spanner on the pipe under the sink’, it will find the right video. You can’t use ChatGPT for that. In other words, the platform understands physicality.”

That claim potentially puts GVTLabs adjacent to one of AI’s next big battlegrounds: Physical AI. This is where AI models need to understand and eventually act within the physical world rather than simply process text, images or abstract instructions.

Robotics companies are increasingly trying to teach machines how objects move, how tools are used, and to be aware of human actions in a space. While GVTLabs is definitely not positioning itself as a robotics startup, the underlying problem it’s tackling is closely related: extracting structured understanding from real-world human behaviour captured on video.

Halstead said its systems ingest not just speech and images, but also “skeletal data” and motion information. “We have skeletal data, motion data,” he said. “There’s a high-value proposition as we train it across billions of hours. But it’s more about the narrative. We then convert all of this knowledge.”

In other words, GVTLabs is trying to understand not merely that a person, wrench and pipe appear in the same frame, but what the person did with the wrench, in what order, and what happened next. That ‘temporal dimension’ is the crux of the technology.

Halstead may have some competition, however. 

Coactive AI, founded in 2021 and having raised $44m in 2024, applies multimodal AI to enterprise image and video libraries. Shofo is a tech startup backed by Y Combinator (YC W26) that claims to be building the world's largest index and dataset library of short-form videos specifically designed for training the next generation of AI and machine learning models. Zyris AI claims to be capable of natural-language search across video, audio and text.

Beyond frame-by-frame AI

However, while most multimodal AI models, such as the ones above, can analyse individual images extracted from video or process a transcript alongside selected frames. Halstead argues that this misses much of what makes video valuable, which is entire events happening through time.

“Other AI video processes just look at a video one frame at a time, but don’t understand what actually happened,” he said. “LLMs don’t have big enough context windows to understand all this video other than by frames.” GVTLabs says it addresses that using what Halstead calls a “temporal agentic harness”.

Rather than ingesting an entire video indiscriminately, the system can decide what it needs to inspect, move through footage and build an understanding of the sequence of events. Halstead demonstrated the approach using an Agatha Christie movie.

“I asked it to look at an Agatha Christie movie but I didn’t let it watch the whole movie,” he said. “It spent five minutes watching the movie and gave the correct answer for the ending, even though it didn’t see the end of the movie.”

Does that mean it will have predictive qualities? Could it predict what happens next on a CTTV camera?

“We don’t do true live video,” says Halstead. “But we can process a one-hour video in, like, 20 seconds. But it’s only like any other AI. It can, if prompted, attempt what a human can when reviewing a video.”

GVTLabs’ immediate business will be a consumer play. Ask a question and the system searches across a creator or media owner’s catalogue and returns a playable clip containing the answer, timestamped and linked to the original footage. You can connect it to your social media accounts like Instagram, as well as YouTube and also upload video from your phone and even create Pinterest-style boards.

Halstead sees it as an alternative to the AI chatbot model. Instead of asking an LLM to synthesise an answer from its training data, AskGVT attempts to find a person who has actually demonstrated or explained the answer on camera. And it can also search YouTube, but not for anything other than keywords.

The company says already has “hundreds of creators” using the product in private beta, according to Halstead. “It downloads all your videos, and you can ask it questions, like ‘what could I do better?’ or do an audit of what you should change in how you present a video.”

Creators will be able to turn their video libraries into public AskPages, where viewers can query everything they have published. They can make those pages free or restrict access behind paid day, weekly or monthly passes.

Halstead compares the model to Substack, but built around accumulated expertise rather than newsletters. “I hope we get millions of people with knowledge starting to use it,” he said. It is, he added, “an expertise play for valuing your knowledge.”

GVTLabs is launching two initial creator products:

AskPage turns a creator’s video library into a public question-and-answer interface, returning clips from their own footage.

Studio gives creators and their teams a private workspace to interrogate their archive, surface clips, analyse patterns and turn findings into reports, decks or other assets.

AskGVT Mobile app

A big raise is on the cards

Halstead told me that GVTLabs raised around £1m from London-based Mercuri VC more than a year ago, but has otherwise been largely bootstrapped by himself.  The company now has around 20 people across the UK and US. Zuzanna Wilson, previously at Huddle and DataSift, is co-founder and CMO.

Halstead says the company expects to raise substantially more capital as it scales. The intended number is not fixed. “The aim is to raise a very large amount,” Halstead said. Referring to AI valuations these days, he said “The number keeps doubling in this wild-west environment.”

UK TV advertising veteran is joining

The potentially larger enterprise opportunity is broadcasters and rights holders sitting on decades of archive footage. To realise this direction, GVTLabs has recruited Rhys McLachlan, formerly ITV’s Director of Advanced Advertising, as Head of Media Partnerships to lead that business.

McLachlan spent six years at ITV, where he built and launched its Planet V addressable advertising platform and created ITV AdLabs. According to figures cited by GVTLabs, Planet V has generated more than £1bn in cumulative advertising revenue and now handles more than 95% of ITV’s digital advertising revenue.

McLachlan will lead three enterprise products: GVTLabs Studio, for searching private video archives; Roster, for finding creators and analysing campaign performance; and Channels, which lets media companies build their own “ask anything” experiences around their content and talent.

For broadcasters, GVTLabs Studio will offer to search vast archives, making their contents queryable. This could allow producers to surface footage faster, create new audience products and eventually make advertising more context-aware.

“What GVTLabs has built is the missing layer for every broadcaster and publisher,” McLachlan said in a statement. “Finally being able to ask your archive a question and get back the exact moment.”

McLachlan previously worked with Halstead through ITV’s relationship with privacy technology company InfoSum, serving as a board observer from 2020 until its sale to WPP in 2025.

A third startup for Halstead

GVTLabs is Halstead’s latest attempt to make a previously difficult form of information queryable.

His earlier startup DataSift helped companies query large volumes of social data, while InfoSum built privacy-preserving technology for joining datasets without exchanging the underlying user data. InfoSum was acquired by WPP in 2025.

But what do they say? Third time’s the charm.

Reply

Avatar

or to participate