T ternura

Hello! I'm

Zixuan Zhu

Currently a core maintainer of Harbor, building agent evaluations with the Terminal-Bench team.

I am an MSc student at the School of Physical and Mathematical Sciences, Nanyang Technological University (NTU), graduating in 2026. I received my dual bachelor's degree from Beijing University of Posts and Telecommunications (BUPT) and Queen Mary University of London (QMUL).

I am looking for PhD, RA, or research-internship opportunities, and I am open to research collaborations. Please feel free to reach out, even just for a coffee chat.

My research currently centers on the following topics:

  • Self-evolving agents: how agents can continually and stably self-improve from their own experience (UCE).
  • Scalable evaluation: how to scale agent evaluation efficiently and in parallel (Harbor), and keep it robust and trustworthy.
  • Economics of AI: how to reduce the token cost of agentic reasoning and get more capability from every token, until capable agents become viable on small, widely accessible models.

I am fortunate to collaborate with Prof. Ludwig Schmidt, Lin Shi, Haowei Lin, Xiaoyue Zhou, and Xiang Li on agent benchmarking, and with Senkang Hu on self-evolving agents.

New blog post: "Introducing Harbor-Index", a compact benchmark for agentic evaluation.

New paper on arXiv: "Unified Context Evolution for LLM Agents".

Unified Context Evolution for LLM Agents

A gradient-free framework that externalizes agent experience into an evolving library of typed, reusable context units, retrieved at decision time and pruned when no longer valuable.

Paper Code (coming soon)

Introducing Harbor-Index

A compact, diverse, challenging, and high-quality benchmark for agentic evaluation.

Blog Dataset

Harbor

A framework for evaluating and optimizing agents and models in container environments.

Website Code

Transformers

The model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Website Code

Loading GitHub contribution data...