Speculative Decoding Speculative decoding accelerates auto-regressive generation in large language models (LLMs) by leveraging a lightweight draft model to predict the next γ tokens. The main LLM then ...
Average decoding scores for modality-agnostic decoders (green), compared to modality-specific decoders trained on data from subjects viewing images (orange) or on data from subjects viewing captions ...
For large base models, you can stream hidden states from a live vllm serve instead of dumping them to disk: a co-located server produces the base-model hidden states on the fly and sends them to the ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results