start here

👋 Hi, I'm

Zimo

Zimo Chen

Business Analytics and Quantitative Finance at NUS. Currently AI Engineer Intern at Temus. Below is a condensed snapshot of my work, involvements, projects, and hackathon entries. Full details on the /work page.

Work

AI Engineer Intern, AI and Data

Aug 2026 to Dec 2026

Temus

  • Currently building Dialogue Forge v2, a real-time AI avatar platform that Temus deploys across client engagements. FastAPI and Pipecat for orchestration, Azure Speech for STT and TTS, AWS Bedrock for the LLM layer, ElevenLabs for voice, and Unreal Engine Pixel Streaming to render the avatar itself.
  • The interesting part is the real-time audio pipeline. Microphone in, Silero VAD for speech segmentation, streaming STT, LLM inference, TTS, all wired so the avatar can handle interruptions and stream a response back without waiting for the full turn to finish. Making that feel like a real conversation instead of a laggy walkie-talkie is most of the challenge.

Real-Time AI Avatar · FastAPI · Pipecat · AWS Bedrock · ElevenLabs · Unreal Pixel Streaming · Python

Risk AI Services Intern

Jun 2026 to Jul 2026

PricewaterhouseCoopers (PwC), Singapore

  • Spent two months on the Risk AI Services team, shadowing project leads across one client engagement from discovery to handoff.
  • Sat in on stakeholder discovery sessions with the business owners. Most of the failure modes I would have called technical a year ago turned out to be scoping problems in disguise.
  • Turned the requirements into an AWS pipeline with AI agents, then helped through implementation and UAT before handoff. I learnt alot translating business needs into technical requirements, and the importance of scoping and communication when doing both the consulting and technical work.

Client Discovery · AWS Cloud Architecture · AI Agents · UAT

Involvements

Teaching Assistant, CS2040 Data Structures and Algorithms

Aug 2026 to Dec 2026

National University of Singapore, School of Computing

Teaching · Data Structures and Algorithms

Derivatives Analyst, Faculty-Advised Student Fund, 200k AUM

Dec 2025 to May 2026

Quantitative Finance Association, UNC Chapel Hill

  • Built option pricing tools across Black-Scholes and binomial trees, stress tested Greeks, and fitted volatility surfaces from historical options data. The final surface cut out of sample RMSE by about 17%.
  • Designed hedge overlays for roughly USD 70k to 90k of equity exposure, with CPI and FOMC stress scenarios to test how the book behaved under two sigma volatility shocks.
  • Built a Python backtester for delta hedged options strategies, including PnL attribution, DV01 sensitivity, and regime diagnostics.

Options Pricing · Greeks · Hedging · Volatility Surface · Python

Research

Shock-to-Equilibrium Forecasting with Generative AI

Aug 2026 to Present

Final Year Dissertation · Supervised by Prof. Huang Ke-Wei

  • Studying how a stochastic process settles after a discrete shock, beginning with the distribution of five-minute stock returns following overnight earnings-call announcements.
  • Curating event-aligned time-series datasets across finance and other domains where recurring, well-defined events create measurable shocks, with an emphasis on many independent series and clean event-to-outcome joins.
  • Benchmarking strong stochastic-process forecasting methods and evaluating a new generative AI approach across multiple datasets, with the broader aim of producing conference-quality empirical evidence.

Generative AI · Time-Series Forecasting · Stochastic Processes · Event-Driven Modelling · Financial Markets · Dataset Curation · Benchmarking

Projects

FinSight: Agentic Financial Research Pipeline

Mar 2026 to Present

Personal Project

  • Built a resumable, evidence-grounded LangGraph pipeline that turns a company, industry, or macro target into an institutional report where every quantitative claim traces to a persisted, approved evidence record.
  • Designed a three-kind evidence contract for web, provider, and analytic records, with DAG-based approval and checkpointed execution that safely resumes interrupted research runs.
  • Developed a deterministic, LLM-free engine for DCF valuation, ratios, forecasts, and peer multiples, paired with a critic that re-derives cited figures and blocks fabricated or unsupported numbers before publication.

Python · LangGraph · DeepSeek · Pydantic · Financial Research · DCF Valuation · Evidence Provenance · SQLite

GitHub ↗

Retrieval-Augmented Generation System

Dec 2025 to Present

Personal Project

  • Built a RAG pipeline from raw PDFs to cited answers because confident answers without sources are not useful in a research workflow.
  • Combined BM25 and dense retrieval with cross encoder reranking, lifting MRR by about 22%.
  • Added citation grounding and low evidence fallbacks, reducing hallucinations by roughly 25% to 35% on a 100 question benchmark.
  • Reached around 70% to 80% grounded answers with correct citations, plus clearer failure modes when the system did not know enough.

Python · FAISS/Pinecone · SBERT/BGE · BM25 · Cross-Encoder Reranking · QwenV3 · FastAPI

Rates Term Structure and Derivatives Analytics Suite

Aug 2025 to Nov 2025

Personal Project

  • Bootstrapped implied SOFR forward curves from futures data and worked through the long end until the curve behaved.
  • Ran swap scenario PnL under parallel and twist shifts, which made the rate exposure much easier to see than duration alone.

SOFR Curves · FRAs · Interest-Rate Swaps · R

View Report ↗·GitHub ↗

Multimodal Video Understanding

Aug 2025 to Feb 2026

Personal Project

  • Built a video classifier that fused frames, audio, and transcripts because the useful context was often split across all three.
  • Compared pooling, temporal attention, late fusion, and learned fusion. Gated fusion performed best after the error analysis made it clear where each stream helped.
  • Gained roughly 8 to 12 macro F1 points over visual only baselines, mostly from audio and transcripts catching what the frames missed.

PyTorch · CLIP/SigLIP · Whisper · OpenCV · LightGBM · Multimodal Fusion · W&B

Multi-Model Market Forecasting and Risk Engine

Feb 2025 to Present

Personal Project

  • Built a macro feature stack with more than 50 inputs, then used PCA denoising to strip out noise instead of adding more complexity. Out of sample information ratio moved from 0.24 to 0.34.
  • Combined ARIMA-GARCH, XGBoost, and LSTM in a walk forward ensemble because markets rarely reward trusting one model too much.
  • Reached 0.78 out of sample Sharpe with 14% max drawdown across more than 10 equities from 2015 to 2024. The Sharpe was useful, but the drawdown was the part that kept the result grounded.

ARIMA-GARCH · XGBoost · LSTM · Walk-Forward Validation · Python

GitHub ↗

Options Volatility and Risk Neutral Distribution Lab

Aug 2024 to Present

Personal Project

  • Scraped options chains and solved for implied volatility, turning raw quotes into a surface I could inspect and test.
  • Derived risk neutral densities with Breeden-Litzenberger and compared the market's implied distribution against realised returns.
  • Validated the densities against realised returns and found the usual fat tails, which made the market's caution look more reasonable than it first seemed.

Implied Volatility · Breeden-Litzenberger · Monte Carlo · Python

GitHub ↗

Hackathons

Multimodal Fraud Detector, 3rd Place

Feb 2026 to Feb 2026

Hacklytics @ Georgia Tech

  • Led a team of four building a fraud risk model from earnings call text, audio, and SEC MD&A. The main challenge was class imbalance, so calibrated LightGBM carried much of the final system.
  • Built call level features through temporal compression and improved results with bootstrap stable late fusion. PR AUC and ticker level confidence intervals helped separate real signal from leaderboard noise.

Python · LightGBM · Sentence-BERT · Time-Series Aggregation

GitHub ↗·Devpost ↗

MoE-Gated Volatility Forecasting, 2nd Place

Dec 2025 to Mar 2026

Martingale Hacks Winter 2025 (Kaggle)

  • Built LightGBM and HAR feature MoE ensembles to forecast next period volatility from lagged market signals under strict out of sample evaluation.
  • Designed leakage safe rolling CV, OOF diagnostics, and ablations across learners and blends. Most of the score gain came from cleaner validation rather than a larger model.
  • Productionized the Kaggle kernels and submission flow with reproducibility checks.

Python · pandas · NumPy · scikit-learn · LightGBM · HAR Features · MoE Ensembles

GitHub ↗·Devpost ↗