Research highlights

Trustworthy language systems, from evidence to impact.

My research asks how language models can find reliable evidence, communicate it at the right level, and remain dependable in high-stakes health and crisis settings.

Grounding & knowledge

Evidence-grounded language models

I develop retrieval-augmented, knowledge-graph, and multi-agent methods that connect model responses to authoritative and timely sources. The goal is not simply to retrieve more text, but to select the evidence needed for a specific claim and make the resulting answer easier to verify.

  • Retrieval-Augmented Generation
  • Knowledge Graphs
  • Multi-Agent Systems
  • Evidence Attribution

Questions I study

  • How should multiple agents retrieve, synthesize, and critique evidence?
  • How can structured knowledge expose missing context and latent information needs?
  • Which evidence is sufficient to support a factual, traceable response?

Selected projects

Published work

People & interaction

Adaptive human–AI interaction

I design systems that adapt explanations to a reader’s literacy, knowledge, intent, and uncertainty. This work treats communication as an interaction: a model may need to adjust explanatory depth, ask a focused question, or answer directly while preserving the same factual foundation.

  • Controllable Generation
  • Health Literacy
  • Personalization
  • Proactive Clarification

Questions I study

  • How can one evidence base support explanations at different levels of detail?
  • When should an AI ask for clarification instead of answering immediately?
  • How can adaptation improve usefulness without changing factual meaning?

Selected projects

Published & ongoing work

Findings of EMNLP 2025

Speaking at the Right Level →

A RAG and reinforcement-learning framework that generates grounded corrections for readers with different health-literacy needs.

Measurement & reliability

Evaluation for high-stakes NLP

I build benchmarks, metrics, and human-evaluation protocols that test whether a system is factual, consistent, useful, and sensitive to context. Health and crisis communication provide demanding test beds because a fluent response can still fail through stale evidence, misplaced information, or inappropriate detail.

  • LLM Evaluation
  • Diagnostic Benchmarks
  • Human Evaluation
  • Crisis Informatics

Questions I study

  • Does a response use evidence from the right place and time?
  • Can automated judges measure factuality and audience alignment reliably?
  • Which metrics reflect what readers actually find understandable and useful?

Selected projects

Published work

Ongoing research

Under review

Under review

Spatiotemporal Reasoning for Crisis Response

An ongoing study of how language models retrieve and reason over location- and time-sensitive information during crises.

A connected agenda

Ground the answer. Adapt the explanation. Test what matters.

These areas reinforce one another: reliable evidence supports adaptation, human needs define meaningful evaluation, and rigorous evaluation reveals where grounding and interaction methods must improve.

See the complete publication record, including papers beyond these selected projects.

View all publications