Open Source

At Pruna AI, we are a model laboratory and inference provider: we create our own performance models, host optimized models on Pruna and with partner platforms, and ship optimization research into production. This section covers our open-source optimization framework pruna for developers who want to compress and optimize their own models.

Before smashing: 4.06s inference time
After smashing: 1.44s inference time

Our compression framework pruna is made by developers for developers. It is designed to make your life easier by providing a seamless integration of state-of-the-art compression algorithms. In just a few lines of code, pruna helps you integrate a range of diverse compression algorithms and evaluate their performance - all in a consistent and easy-to-use interface.

Pruna Open Source

pruna is a free and open-source compression framework that allows you to compress and evaluate your models.

Install Pruna

Learn how to install pruna and use serving integrations.

Install Pruna
Smash your first model

Understand how to use pruna to compress and evaluate your models.

Smash your first model
Evaluate and benchmark your models

Learn how to benchmark and evaluate your optimized models with pruna.

Evaluate quality with the Evaluation Agent
Tutorials

Get familiar with end-to-end examples for various specific modalities and use cases.

Tutorials Pruna

How does it work? First, you need to install pruna:

pip install pruna

After installing pruna, you can start smashing your models in 4 easy steps:

  1. Load a pretrained model

  2. Create a SmashConfig

  3. Apply optimizations with the smash function

  4. Run inference with the optimized model

Let’s see how it works with an example:

import torch
from diffusers import StableDiffusionPipeline
from pruna import smash, SmashConfig

# Define the model you want to smash
pipe = StableDiffusionPipeline.from_pretrained(
    "CompVis/stable-diffusion-v1-4",
    torch_dtype=torch.float16
)
pipe = pipe.to("cuda")

# Initialize the SmashConfig
smash_config = SmashConfig(['stable_fast', 'deepcache'])

# Smash the model
smashed_model = smash(
    model=pipe,
    smash_config=smash_config,
)

# Run the model on a prompt
prompt = "a photo of an astronaut riding a horse on mars"
image = smashed_model(prompt).images[0]

Now that you’ve seen what pruna can do, it’s your turn!

Pruna Community

We love to organize events and workshops and there are many coming up! Find more info about our community and events in the Community section.