AnyWit Robotics deeply engages in the field of embodied intelligence and is committed to innovating human-robot interaction methods. We focus on providing hardware and software solutions for multimodal emotional interaction robots, aiming to create robots with "vitality"
Research and develop multimodal generative models to achieve breakthroughs in human–robot interaction experience and computational efficiency.
Job Responsibilities
Research and advance the application of multimodal models in face-to-face interactive robots; explore and build innovative algorithm models and technical solutions.
Deeply participate in data construction, and design, develop and optimize multimodal generative models.
Keep track of cutting-edge advances in multimodal AI technologies, and conduct technical research and summary.
Job Requirements
Master’s degree or above in Artificial Intelligence, Big Data, Computer Science or related fields. First-author publications in CCF-A journals/conferences or equivalent venues on topics including speech synthesis, image/video generation, VLM, audio large models, etc.
Proficient in architecture design and training optimization of multimodal generative models (e.g., VLM, audio large models, video generation models); solid command of mathematical principles and training strategies for cutting-edge frameworks such as Transformers and diffusion models.
Familiar with multimodal dataset construction workflows and quality standards; capable of developing scripting tools for dataset construction and data cleaning.
Master systematic model evaluation methodologies and inference optimization techniques.
Experienced in complete research projects; able to clearly elaborate scientific problems, research approaches, technical routes and core conclusions in a structured manner.
Passionate about human–robot affective interaction, with in-depth understanding of edge-cloud collaboration mechanisms.
Strong logical thinking and capable of designing rigorous and reasonable experimental schemes.