Skip to content
 
 

Repository files navigation

SentencePiece Parallel

This is not an official Google product.

Technical highlights

  • This fork utilizes multiple cores and smart buffering to divide-and-conquer large datasets. Takes a 10B token source from over an hour to under 10 minutes. Reduces memory overhead by ~50 percent.

About

Fork of Google's SentencePiece tokenizer that adds better memory management for large corpus volumes and increased throughput by parallelizing various parts.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages