RT-DETR TransformerOpen VocabularyVisual EmbeddingsLock ModeReal-Time 35+ FPSFastAPI + ONNXResNet-18 FeaturesMulti-Pipeline FusionLearn New Objects80 COCO ClassesYOLO-WorldCombined DetectionRT-DETR TransformerOpen VocabularyVisual EmbeddingsLock ModeReal-Time 35+ FPSFastAPI + ONNXResNet-18 FeaturesMulti-Pipeline FusionLearn New Objects80 COCO ClassesYOLO-WorldCombined Detection
Three Detection Pipelines
Pick Your Detection Mode
01DETECTION
KNOWN OBJECTS
RT-DETR-L
Pre-trained 80 COCO classes detected with transformer-based real-time architecture.
Launch
02DISCOVERY
OPEN VOCABULARY
YOLO-World
Discover any object from natural language descriptions. Zero-shot detection at 35+ FPS.
Launch
03RECOGNITION
VISUAL EMBEDDINGS
ResNet-18
Teach the system new objects and recognize them later through visual similarity matching.
Launch
How It Works
AI Processing Pipeline
01
CAPTURE
Webcam or image upload
02
DETECT
Multi-pipeline inference
03
RECOGNIZE
Open vocabulary + visual embeddings
04
ANALYZE
Classify, embed, compare
05
TRACK
Lock Mode persistent targeting
06
RESULTS
Annotated live output
LOCK MODE
Persistent Target Tracking
Lock onto any object and maintain persistent tracking across video frames. Combined pipeline fuses RT-DETR, YOLO-World, and visual embeddings for robust identification.