Skip to content

Installation

This guide will help you get TorchLingo installed and ready to use in just a few minutes.

The fastest way to get started! Google Colab requires zero setup on your machine and provides free GPU access.

Step 1: Open a Colab Notebook

  1. Go to Google Colab
  2. Click File → New notebook (or open one of our tutorial notebooks)

Step 2: Enable GPU (Important for Training!)

  1. Click Runtime → Change runtime type
  2. Select GPU from the "Hardware accelerator" dropdown
  3. Click Save

Free GPU Access

Colab provides free access to NVIDIA GPUs. Training is 10-50x faster with GPU compared to CPU!

Step 3: Install TorchLingo

In the first cell of your notebook, run:

%pip install torchlingo

Step 4: Verify Installation

# Check that everything is working
import torch
print(f"PyTorch version: {torch.__version__}")
print(f"CUDA available: {torch.cuda.is_available()}")
if torch.cuda.is_available():
    print(f"GPU: {torch.cuda.get_device_name(0)}")

# Test TorchLingo
import torchlingo
from torchlingo.config import get_default_config
config = get_default_config()
print(f"\n✓ TorchLingo is ready! Default batch size: {config.batch_size}")

You should see output showing CUDA is available and a GPU name (like "Tesla T4").


💻 Local Installation

If you prefer to run locally or want to explore/modify the source code:

Prerequisites

Before installing TorchLingo, make sure you have:

  • Python 3.10 or higher - Download Python
  • pip (comes with Python)
  • Git (optional, for cloning the repository)

Check your Python version

Open a terminal and run:

python --version
You should see Python 3.10.x or higher.

Method 1: Install with pip (Simplest)

pip install torchlingo

Method 2: Install from Source (For Development)

This is the best method if you want to explore the code, run tutorials, or modify things:

# Clone the repository
git clone https://github.com/byu-matrix-lab/torchlingo.git
cd torchlingo

# Create a virtual environment
python -m venv .venv
source .venv/bin/activate

# Install in editable mode
pip install -e .
# Clone the repository
git clone https://github.com/byu-matrix-lab/torchlingo.git
cd torchlingo

# Create a virtual environment
python -m venv .venv
.venv\Scripts\activate

# Install in editable mode
pip install -e .

Install Git LFS before you clone

Two files are too large for ordinary version control and are stored in Git LFS instead: the example corpus (data/example.tsv) and the pretrained checkpoint (data/pretrained/model.pt).

Without Git LFS, git clone still succeeds, but those two files arrive as three lines of text describing the real file rather than the file itself. Tutorials 2 through 5 then fail with errors that look like bugs in the library and are not.

Install it once, before cloning:

git lfs install

If you already cloned without it, you do not need to start over:

git lfs install
git lfs pull

To check, look at the size of the corpus. It should be about 17 MB, not 130 bytes.

Verify Installation

Let's make sure everything is working:

# In a Python shell or script
import torchlingo
from torchlingo.config import get_default_config

config = get_default_config()
print(f"TorchLingo is ready! Default batch size: {config.batch_size}")

You should see output like:

TorchLingo is ready! Default batch size: 64

Installing Dependencies

TorchLingo has a few key dependencies that are automatically installed:

Package Purpose
torch Deep learning framework
pandas Data loading and manipulation
sentencepiece Subword tokenization
tensorboard Training visualization
sacrebleu Translation quality metrics

Optional: Language-Specific Dependencies

TorchLingo supports specialized tokenizers for Asian languages. Install based on your needs:

pip install torchlingo[japanese]

Installs fugashi and unidic-lite for Japanese morphological analysis.

pip install torchlingo[chinese]

Installs jieba for Chinese word segmentation.

pip install torchlingo[asian]

Installs tokenizers for both Japanese and Chinese.

Extra Packages Installed Use Case
japanese fugashi, unidic-lite Japanese tokenization
chinese jieba Chinese word segmentation
asian fugashi, unidic-lite, jieba Both Japanese and Chinese support

Optional: Development Dependencies

If you want to run tests, lint code, or contribute to TorchLingo:

pip install -e ".[dev]"

This installs:

Package Purpose
ruff Fast Python linter and formatter
isort Import sorting
build Building distribution packages
twine Publishing to PyPI

Optional: Documentation Dependencies

To build these docs locally:

pip install -e ".[docs]"

GPU Support

TorchLingo works on CPU out of the box, but training is much faster with a GPU.

PyTorch should automatically detect your NVIDIA GPU. Verify with:

import torch
print(f"CUDA available: {torch.cuda.is_available()}")
print(f"GPU: {torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'None'}")

On M1/M2/M3 Macs, PyTorch can use the Metal Performance Shaders backend:

import torch
print(f"MPS available: {torch.backends.mps.is_available()}")

No GPU? No problem! All tutorials work on CPU. Just expect longer training times for larger models.

Common Issues

ImportError: No module named 'torchlingo'

Make sure you've activated your virtual environment:

source .venv/bin/activate  # macOS/Linux
.venv\Scripts\activate     # Windows

ModuleNotFoundError: No module named 'torch'

PyTorch didn't install correctly. Try:

pip install torch --upgrade

Permission denied when installing

Don't use sudo pip install. Use a virtual environment instead (Method 1 above).

Next Steps

Now that TorchLingo is installed, let's build something!

Quick Start