Updated
Updated · arxiv.org · Aug 28
DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging
Updated
Updated · arxiv.org · Aug 28

DARTS: Decoder-Aware Representation Tuning via Surgery for Model Merging

1 articles · Updated · arxiv.org · Aug 28

Summary

  • Researchers have introduced DARTS, a new method to improve model merging for large language models by addressing representation bias in decoder architectures.
  • DARTS uses entropy-weighted loss and position-dependent corrections, achieving significant performance gains across code, math, and instruction-following tasks with minimal added parameters.
  • This approach may streamline multi-task model deployment and closes the performance gap between merged models and individual fine-tuned models, especially in autoregressive settings.