Haoyue Tang–Research Interests

My research aims at exploring the fundamental limits, designing efficient algorithms and prototyping for decision making in stochastic networks and stochastic systems.

Recently I am interested in:

  • Efficient Long Sequence Modeling through Software-Hardware Co-Design

  • Foundations for Diffusion and Flow Matching Models

  • Applications of Machine Learning Methods in Healthcare and Networking

Below are my past research projects:

Classifier-Free Guidance in Diffusion Models

dpg 

Classifier-free guidance (CFG) is the standard mechanism for conditional sampling in diffusion models, and essentially every deployed text-to-image system depends on it. The goal of this project is to understand the mechanism of CFG, develop efficient ways to implement CFG in inverse problems, and improve the CFG method in general.

Publications

Grants

  • NAIRR260411: GPU Resources for Teaching Information-Theoretic Deep Generative Models, ECE 693B, University of Hawaiʻi at Mānoa

  • NSF ELE260074: Diagnosing and Improving Conditional Score Estimation for Classifier-Free Guidance

Efficient Long Sequence Modeling through Software-Hardware Co-Design

sequencemodeling 

Efficient long sequence modeling architecture is important for applications such as recommendation systems and drug discovery. For example, in recommendation systems, a sequence representing the past user interaction logs is first summarized into a few tokens, and then fed into the deep neural network for further user-content interaction prediction. Utilizing the attention structure for sequence modeling faces two challenges: (1) summarizing the full long behaviour sequence is GPU-memory costly; (2) over long sequences the attention weight is also over-allocated to irrelevant events, reducing the accuracy of user interest modeling. How to understand and model long sequences efficiently and accurately remains challenging.

Our research attempts to overcome the aforementioned challenge through a semantic-aware, hierarchical sequence pruning mechanism developed jointly with the attention kernel that executes it. Unlike prior approaches that use the full sequence for modeling user interests, we propose a synergistic software-hardware co-design approach for hierarchical long sequence pruning and modeling.

Grants

  • NAIRR260485: Hierarchical Semantic-Aware Sequence Pruning for Efficient Long-Context User Modeling

Online Learning for Network Optimization

1350 

How to optimally manage the freshness of information updates (have a good estimation about \(X_t\) at the destination) sent from a source node to a destination via a channel when the channel statistics is unknown? By using the Age of Information (AoI) as a freshness metric, we first present a stochastic approximation algorithm that can learn the optimum sampling policy almost surely, and prove that the cumulative regret of the proposed algorithm is minimax order optimum. By incorportating more information on the time-varying process and design content agnostic data collection policies, our algorithm can lower the estimation error of \(X_t\). This project is supported by NSF-AI Institute Athena. Slides, Poster

Highlights

  • Theoretic Contributions:

    • New convergence results for stochastic optimization in an open set.

    • New converse results for online learning algorithms based on non-parametric statistics.

Publications

Cross-Layer Scheduling for Data Freshness Optimization

140 

Previous work reveal that, to keep data fresh, it is important to guarantee: (i) low latency; (ii) high data rate; and (iii) service regularity. Considering sensors in wireless networks have energy constraints and the wireless channels are time-varying, how to opportunistically generate, transmit and deliver data so that the the multi-objective optimization problem can be settled? Based on dynamic programming, bandits and large deviation analysis, I propose a joint data sampling, power control and scheduling algorithm that is optimum in large scale networks.

Highlights

  • The first optimal multi-user scheduling algorithm in AoI literatures: we show that for a network with \(N\) users and \(M\) bandwidth, by fixing \(N/M\) as a constant, the average AoI optimality gap between the proposed algorithm and the lower bound is \(\mathcal{O}(1/\sqrt{N})\), indicating that the proposed algorithm is optimal in large-scale networks. ITW2020 presentations Allerton 2019 presentations

Publications

Signal Processing for mmWave Channel Estimation

By exploiting spatial sparse structure in mmWave channels, we propose an angle domain off-grid channel estimation algorithm for the uplink millimeter wave (mmWave) massive multiple-input and multiple-output (MIMO) systems. The proposed method is capable of identifying the angles and gains of the scatterer paths. Comparing the conventional channel estimation methods for mmWave systems, the proposed method achieves better performance in terms of mean square error. Numerical simulation results are provided to verify the superiority of the proposed algorithm.

Publications

Domain Adaptation and Out-of-Distribution Generation using Causal Inference

350 

Conventional supervised learning methods, especially deep ones, are found to be sensitive to out-of-distribution (OOD) examples, largely because the learned representation mixes the semantic factor with the variation factor due to their domain-specific correlation, while only the semantic factor causes the output. To address the problem, we propose a Causal Semantic Generative model (CSG) based on a causal reasoning so that the two factors are modeled separately, and developed a variational Bayesian method for training CSG, a method for OOD prediction from a single training domain. This work is done during my internship at Microsoft Research. Poster

500 
  • We prove that under certain conditions, CSG can identify the semantic factor by fitting training data.

  • Empirical study shows improved OOD performance over prevailing baselines.

My publications and presentations contains more information about my past and present research.