Yongjun “Teo” Choi
From Teo Wiki, a public encyclopedia-style profile page
Yongjun “Teo” Choi is an M.S. student at UNIST in the 3D Vision & Robotics Lab, advised by Prof. Kyungdon Joo. His work sits at the intersection of computer vision, generative visual editing, scene understanding, video understanding, visual-language modeling, and audio-visual learning.
Teo’s current public research direction can be summarized as building practical multimodal AI systems: models and applications that understand, edit, compare, and reason over real-world visual and audio-visual signals. His portfolio includes diffusion-based image manipulation, reference-guided video anomaly detection, human–robot interaction, and language-guided 3D scene understanding.
Ask about Teo
This page is designed for collaborators, professors, lab members, and international readers who want a quick map of Teo’s background, research taste, and public work. The search box and cards work as section navigation. Start with one of these questions:
Overview
Teo Wiki is a public, encyclopedia-style companion to Teo’s academic portfolio. It is intentionally more direct and context-rich than a one-page CV: its purpose is to help readers quickly understand the person behind the project list—his research identity, technical taste, public projects, and near-term questions.
The official portfolio remains the best place to see publications, CV files, project links, and formal achievements. This wiki adds interpretation: how the pieces connect and what kind of collaborator Teo is likely to be.
Following Andrej Karpathy’s “LLM Wiki” pattern, Teo Wiki is intended to grow in two layers: this public page as the readable front door, and a private markdown research wiki where sources, notes, concepts, and questions compound over time.
Research identity
Teo’s research has focused on 3D scene understanding, image manipulation using generative models, video understanding, and visual-language modeling. More recently, he has been expanding toward audio-visual modeling, with an emphasis on developing and evaluating multimodal learning frameworks.
- Visual editing: realistic image manipulation with generative models, including diffusion-based hair removal.
- Scene and 3D understanding: object-centric and language-guided reasoning over visual scenes.
- Video understanding: comparing, aligning, and detecting anomalies in real-world video streams.
- Audio-visual learning: multimodal modeling that connects sound, vision, and temporal structure.
- Practical AI systems: prototypes and applications that work under real constraints, from industrial inspection to robot interaction.
Selected projects
Realistic Hair Removal / AnyBald
A mask-free diffusion-based framework for realistic hair removal in the wild, designed to preserve facial identity while producing natural bald manipulation.
Reference-Guided Video Anomaly Detection
A practical video anomaly detection system for automated display inspection, using alignment and comparison to identify defects during long-term multi-device operation.
Vision-Based Gomoku AI Robot
A low-cost human–robot interaction system for strategic board games, combining real-time vision perception, reinforcement-learning-based decisions, and robotic arm control.
Lang-Grouping
An object-centric semantic grouping project for language-guided 3D scene understanding.
Cheese! Generative Editing
A generative image editing application project exploring how modern editing models can become practical creative tools.
Research brain and LLM Wiki
Teo Wiki is designed to become more than a static profile. The intended long-term pattern is a curated public surface backed by a private, LLM-maintained research wiki: raw sources stay immutable, while concepts, project notes, comparisons, and valuable question-answer sessions are compiled into interlinked markdown pages.
Public Teo Wiki
The page visitors read: identity, research map, projects, open questions, and public-safe summaries.
Raw sources
Papers, project notes, reading notes, transcripts, diagrams, and drafts kept as source-of-truth material.
LLM-maintained markdown graph
Entity, concept, project, comparison, and query pages with wikilinks, provenance, and update logs.
Publications
- AnyBald: Toward Realistic Diffusion-Based Hair Removal In-the-Wild. Yongjun Choi*, Seungoh Han*, Soomin Kim, Sumin Son, Mohsen Rohani, Edgar Maucourant, Dongbo Min, Kyungdon Joo. Accepted, WACV 2026
- RG-VAD: Reference-Guided Video Anomaly Detection for Automated Display Inspection. Yongjun Choi, Gyeongsu Cho, Jinhyeok Kim, Changsu Ha, Sanggyu Biern, Kyungdon Joo. Under review
- Demonstrating a Vision-Based AI Robot for Strategic Board Games. Taehwan Kim*, Dokeun Lee*, Seonghyeon Kim*, Yongjun Choi*, Sungjun Heo, Thi Thuy Ngan Duong, Kyungdon Joo, Namhun Kim, Jeong hwan Jeon, Hyemin Ahn. Technical report
Experience
Visiting Student · University of Toronto
Jan 2025 – Jun 2025. Special MEng student at MIE, with graduate-level coursework and an industrial project with Modiface.
Research Assistant · 3D Vision & Robotics Lab, UNIST
Mar 2024 – Aug 2026. Research on video understanding and language-guided 3D scene understanding, plus industrial collaboration and teaching assistant work.
Software Developer Intern · Upsight
Jun 2023 – Aug 2023. Contributed to building detection and mobile community app development work.
Undergraduate Research Intern · Computer Vision Lab, University of Seoul
Apr 2022 – Jul 2023. Studied Masked Image Modeling-based ViT and CAM methods; built AI-based feature extraction for an urban forest management mobile platform.
Values and operating principles
- Build things that can be tested. Teo’s public work often connects research ideas to working prototypes, demos, or applied systems.
- Respect both models and data. Good AI systems need strong modeling choices, careful evaluation, and realistic data assumptions.
- Bridge research and application. The most interesting problems are often found where academic methods meet deployment constraints.
- Stay multimodal. Vision, language, audio, and time should be treated as connected signals rather than isolated modalities.
Open questions
- How can generative visual editing systems remain identity-preserving, controllable, and robust in the wild?
- How can models compare long video streams reliably enough for industrial inspection and real-world monitoring?
- What evaluation protocols make audio-visual learning trustworthy beyond benchmark accuracy?
- How can 3D scene understanding and multimodal reasoning become useful in practical interactive systems?
Public/private boundary
This page is intentionally public-safe. It summarizes information that is already suitable for a public academic website: affiliations, projects, publications, research interests, experience, and broad technical directions.
It does not include private personal data, private relationship information, private correspondence, raw notes, confidential industrial details, unpublished experimental specifics, or anything that would be inappropriate to share with a general public audience.