Skip to content
 
 

Repository files navigation

LoongServe

This is an implementation of the paper: "LoongServe: Efficiently Serving Long-context Large Language Models with Elastic Sequence Parallelism".

To reproduce all the main results in our paper, please check the artifact folder and follow the instructions in it.

About

No description, website, or topics provided.

Resources

Stars

Watchers

Forks

Releases

Packages

Contributors

Languages