TrustAndSafety

PYLER

VideoUnderstanding

Tech

Content Moderation

Multi-modal AI

Finding the One Harmful Second in Hours of Video

2026. 7. 19.

Pyler Video Safety featured image

Inside the video safety AI that won NVIDIA’s hackathon and joined NVIDIA’s official case studies

A cooking video you put on while sitting with your kid. Vegetables get chopped, a pot comes to a boil, everyone laughs and passes plates around. Then, somewhere in the middle, a caption flashes in the corner of the screen for a second, and it’s explicit. Or the song in the background turns out to be a piece of hate speech set to music.

A harmful video is rarely harmful from start to finish. Usually it’s an ordinary clip with violence, hate, or a doctored scene slipped into a few moments of it. That’s what a human reviewer misses, and what an AI that grades the whole video at once misses too. PYLER set out to solve that exact problem: find the harmful moment, not just label the file.

1. In the Age of Infinite Video, Who Keeps Us Safe?

Consider what happened recently. One of the best-known names in generative AI abruptly pulled the video generation service it had launched only months earlier. The service had been popular enough to top the app store charts, which made the sudden shutdown all the more startling to the industry. Several reasons were cited officially, but the controversy over deepfakes and harmful generated content is widely understood to have weighed heavily in the decision. The lesson was hard to miss: without a safeguard that catches generated output before it reaches users, even a world-leading company can find itself with no choice but to shut the service down.

Video now accounts for more than 80% of global internet traffic, and the volume uploaded every day continues to climb. Somewhere in that flood, however, sits content nobody should have to encounter: hate speech, graphic violence, sexual material, illegal drugs, deepfake-driven misinformation.

For the past decade, the industry’s answer to this problem has been people. Platforms hired tens of thousands of content moderators and asked them to watch the worst of the internet, all day, every day. We know how that story ends. Moderators develop PTSD and lasting psychological trauma from the job itself. And even setting the human cost aside, manual review has a structural flaw: judgments shift with a reviewer’s cultural background and personal standards, making it nearly impossible to ensure consistent safeguarding across a platform. Add generative video AI to the mix, and this decade-old approach is pushed past its limits, no longer a workable line of defense. The shutdown described above was exactly that signal.

We believe reviewing harmful video should no longer depend on human sacrifice, and in truth, it is no longer a problem human effort alone can solve. The alternative is an AI that verifies content before it circulates, applying standards that are transparent, consistent, and independent of any single reviewer’s judgment. That is the vision behind PYLER’s Trust & Safety research, and we believe the industry can no longer afford to treat it as optional.

2. The Research Problem: Finding the Exact Moment, Not Just the Video

Most safety AI today works on text or single images. Even video systems regularly make only one judgment for the entire file: safe or unsafe. Real moderation needs more. It needs to know exactly when harm appears and what kind it is, down to the second, at what we call the chunk level.

This is a genuinely difficult multimodal problem, and PYLER’s AI Research team has been building toward it by drawing on recent work in the field.

Training data and infrastructure built for the task

We train on safety benchmark datasets that reflect current global standards, covering categories from sexual content and violence to misinformation, illegal activity, and hate speech, across large volumes of both real and synthetic video. Processing video at this scale without bottlenecks requires investment, so we built a computing setup that flexibly draws on high-performance GPU clusters as workload demands change.

A reinforcement learning pipeline that the ecosystem didn’t offer

Our models build on proven open-source vision-language models (VLMs), which we fine-tune extensively for the safety domain. The harder part was temporal grounding: teaching a model to follow the flow of time in a video and point to the specific segment where harm appears. The open-source ecosystem did not support this, so we designed and built our own reinforcement learning (RL) pipeline.

Evaluation needed the same care. We combined several proven open-source techniques into our own video cut-detection pipeline and built evaluation sets in which harmful segments were deliberately spliced into otherwise ordinary footage. Against existing baselines, our model achieved substantial gains in both F1 score and TIoU (Temporal Intersection over Union), a metric that measures how accurately a model predicts a specific timeline.

One qualitative test made the difference clear. We took a normal sports broadcast and inserted a few seconds of gunfire, both sound and footage, into the middle. Baseline models missed it because they read the video as a whole. Our model flagged exactly those seconds, and nothing else.

Just as important, the entire application runs as a complete standalone solution. It neither sends data to nor depends on any external generative platform. For enterprise situations where data leakage is a non-starter, this architecture is a core requirement, not a side benefit.

3. Validation from the Global AI Ecosystem

Research claims are easy to make, so external validation matters to us.

Recently, PYLER won the Domain-Specific Model track at NVIDIA’s Nemotron hackathon (Dev Day Seoul) with this video safety technology. Within the event’s tight time limit, the team built its reinforcement learning pipeline and carried the model through inference, a result the judges singled out. NVIDIA has also published PYLER’s video analysis work as an official case study on its website, a separate recognition that we take as a signal that this technology has moved beyond one startup’s internal project and into the reference set for how the industry discusses video safety.

Read More: NVIDIA Official Case Study — PYLER

The conversation is continuing at the research level as well. We are working with Content Safety researchers at leading global AI companies on one of the field’s open design problems: how to bridge the gap between abstract safety policies and the concrete actions that actually appear in a video, through better rubric and taxonomy design.

4. Where the Research Meets the Market: Why Safety in Advertising Is Trust & Safety

An AI meant to screen the world’s video cannot be completed inside a lab. It needs real video and immediate feedback, and PYLER finds both in the global brand advertising market. Our Brand Safety products already protect top-tier clients in Korea and around the world, and the next step is scene-level targeting and analysis, which breaks a single video into hundreds or thousands of metadata points to find the exact scene a brand is looking for.

And safety in the advertising market is not a world apart from Trust & Safety. A great deal of harmful video exists online because it earns ad revenue. When ads stop flowing to harmful content, its creators stop making money, and when the money stops, so does the incentive to produce more. The more widely video guardrail layers are deployed across the ecosystem, the more harmful content shrinks not at the distribution stage but at the point of creation. What PYLER does in the ad market is cut off the revenue that funds harmful video, and that is the most practical contribution we can make toward the moderation infrastructure described above.

The Connected TV (CTV) market, where PYLER is focused today, shows most clearly the bar this technology must clear. To actually work in this market, a model has to quickly and precisely read the moments in a video and turn them into metadata. It has to understand not just speech but the imagery, logos, and motion in the frame, and serve an ad at exactly the right contextual moment, even on live broadcasts. It also has to judge the safety of incoming ad creatives before they are served, so that an ad meant for adults never plays while a parent and young child are watching TV together. PYLER is close to completing technology at the level the CTV market demands and is fully committed to bringing it into real-world operation. And this is the level video understanding AI must reach before it can be put to serious use in content moderation.

Every capability proven in advertising today feeds directly into tomorrow’s moderation infrastructure. In that sense, commercial success and the Trust & Safety mission are not two separate tracks. They are a single road leading to the same destination.

5. Conclusion: Toward a Consistent, Safe Video Ecosystem

PYLER’s solution is proving itself today on the front lines of brand safety, but its ambition is larger than advertising. What we are building is an integrated verification and moderation layer, one designed to sit at the center of video platforms everywhere and check content before it ever reaches an audience.

The era in which digital safety was bought at the cost of human suffering should end. And with generative video models now enabling the creation and editing of video from a single prompt, platforms and viewers are exposed to harmful content more easily than ever before. That is exactly why an independent AI must stand in its place, filtering the world’s video against standards that stay consistent and transparent, no matter who is watching or where. That is the future PYLER is building toward, and we intend to set the standard for a video ecosystem everyone can trust.

© 2026 PYLER. All rights reserved.

pylerbiz@pyler.tech

© 2026 PYLER. All rights reserved.

pylerbiz@pyler.tech