Hi, I'm

Zhishuo

Ph.D. candidate in Computer Science, focusing on multimodal learning, affective computing, and robust speech understanding.

NowCurrently at PolyU · Hong Kong from Sep 2026

About

I am a Ph.D. candidate in Computer Science at Sichuan University. My research focuses on multimodal learning, affective computing, and robust speech understanding, with particular interests in multimodal sentiment analysis, cross-modal contrastive optimization, noise-resilient speech recognition, and emotion-aware agent workflows. I care about research that remains interpretable and generalizes under real-world conditions.

Research Focus

Core Research

Multimodal Emotion Understanding

Modeling emotion from speech, vision, and text with depth-aware representations.

Systems

LLM + MoE Systems

Task-adaptive routing and efficient expert collaboration for better generalization.

Impact

Applied AI Products

Bridging research and deployment through reliable workflows and automation.

Selected Projects

work in progress
Research:

HEME

Hierarchical emotion modeling with adaptive multi-level mixture-of-experts.

Platform:

Emotion Agent Stack

End-to-end pipeline for multimodal emotion analysis and conversational AI.

Workflow:

AV-RISE

Hierarchical cross-modal denoising for robust audio-visual speech representation under noisy real-world conditions.

Tooling:

Paper-to-Product Toolkit

Toolchain for transforming research prototypes into reproducible demos.

News

from the notebook

Blog

Research Notes Long-form notes on research systems, robust speech, and emotion intelligence.