Software Engineer, Perception & Vision
Laelaps
Zürich · Switzerland
Published : Sep 27, 2026
Every opportunity links to its original source. No employment guarantee.
About this opportunity
Our Mission At Laelaps AI, we believe robotics is entering a transformative decade, much like the arrival of the internet. Advances in AI, cloud computing, and hardware are reshaping what autonomous systems can do. Our mission is to build the intelligent software that powers physical security in the real world - enabling robots and sensors to handle…
Read the full description
Our Mission At Laelaps AI, we believe robotics is entering a transformative decade, much like the arrival of the internet. Advances in AI, cloud computing, and hardware are reshaping what autonomous systems can do. Our mission is to build the intelligent software that powers physical security in the real world - enabling robots and sensors to handle dangerous and critical tasks that humans shouldn't have to.
By engineering the orchestration layer for intelligent security, we aim to create a world that is safer, more secure, and more resilient. We're a strong founding team based in Zurich, backed by visionary investors and advisors. We are engineering the future of security today! The Role As a Computer Vision Engineer, you will build and own the perception systems behind our multi-sensor surveillance and robotics stack.
This is a builder's role first: you start from the best available models, off-the-shelf where they do the job and trained in-house where they don't, and turn them into reliable, real-time systems that detect, track, and re-identify objects across cameras and run efficiently on edge hardware. We care less about novel papers and more about whether what you build works on a live site and stays working.
What You'll Work On Design and ship cross-camera object re-identification (ReID) for people and vehicles across distributed camera networks, in all light and environmental conditions. Build high-throughput object detection and multi-object tracking pipelines that run across many simultaneous video streams. Integrate and adapt vision-language models (VLMs) for open-vocabulary detection, scene understanding, and operator-facing situational awareness.
Own the video ingestion and streaming path (RTSP, WebRTC) from camera to model, with attention to latency, resilience, and dropped-frame handling. Optimize and deploy models on edge hardware: TensorRT, quantization (INT8/FP16), pruning, and other techniques to hit real-time targets on Jetson and edge-class devices. Evaluate, fine-tune, and integrate existing open-source and commercial models, and train custom models when off-the-shelf options fall short, knowing when each is the right call.
Work closely with the software and robotics teams so perception output feeds downstream autonomy and alerting. Who We're Looking For We're looking for a strong engineer who has shipped computer vision or ML systems into production, ideally in an embodied, multi-camera, or multi-modal setting. You think rigorously a