# Getting Started

> **Note**: This documentation is a **work in progress** and subject to further improvements.

First, follow **[Installation Guide](installation.md)** to install WaveFactor into a virtual environment and activate said environment.

## Example Code Snippet

This example shows how WaveFactor can be applied on data and how posterior estimates can be extracted.

```python
import wavefactor as wf

# 1. Simulating input data
# Input requirements:
#   X: continuous expression matrix (N_spots x N_genes)
#   coords: (N_spots x 2) integer grid coordinates covering [0, L-1] x [0, L-1]
#   N_spots = L * L must be an exact power of 4 (e.g. 64, 256, 1024, 4096)
np.random.seed(42)
coords = np.mgrid[0:32, 0:32].reshape(2, -1).T
X = np.random.randn(1024, 200)

# 2. Instantiate WaveFactor estimator
model = wf.WaveFactor(
    n_factors=10,               # Number of latent factors (K)
    n_length_scales=4,          # Wavelet detail levels (R = 5 total resolutions)
    n_init=5,                   # Multi-start initializations (selects best ELBO)
    n_jobs=1,                   # Parallel workers across initializations
    random_state=42,
)

# 3. Apply WaveFactor to the data
factors = model.fit_transform(X, coords)

# 4. Access posterior estimates
result = model.get_result()
print(f"Final ELBO: {result.elbo:.2f} across {result.n_iter} iterations")

spot_factors = result.factors        # Latent spatial factors (N_spots x K)
gene_loadings = result.loadings      # Factor loadings matrix (K x N_genes)
gene_pips = result.gene_pip          # Gene Posterior Inclusion Probabilities
spatial_pips = result.spatial_pip    # Multiresolution spatial wavelet PIPs
```

## Another Example With Simulated Data Based On Gene Program Ground Truths

WaveFactor includes additionally ready-to-run simulation example with ground truth gene programs, fits the model, and checks the results against the ground truth gene programs.

This example can be ran with the following command.

```bash
python examples/getting_started/run_analysis.py
```

---

# WaveFactor Documentation

!!! info "Work in Progress"
    This documentation is a **work in progress** and subject to further improvements.

**WaveFactor** (formerly `WaviFM`) is a Bayesian factor modeling framework for spatial transcriptomics that explicitly models spatial length scales by performing Coordinate Ascent Variational Inference (CAVI) directly on 2D Discrete Wavelet Transform (DWT) coefficients.

---

## 📚 Contents

- **[Installation Guide](installation.md)**: How to set up an isolated environment and install WaveFactor.
- **[Data Input & Output](input_output.md)**: Expected input grid formats, preprocessing tips, and model outputs.
- **[Getting Started Example](getting_started.md)**: How to run the quick self-contained simulation.

> 🤖 **LLM Context**: Working with an AI assistant? Use the raw single-file documentation at [llms.md](llms.md).


---

# Data Input & Output

> **Note**: This documentation is a **work in progress** and subject to further improvements.

---

## 1. What WaveFactor Expects as Input

WaveFactor needs two inputs: an expression matrix `X` and spot coordinates `coords`.

### A. The Spatial Grid
- **Square Grid**: The data must sit on a square grid where each side length $L$ is a power of 2 (e.g., $16 \times 16$, $32 \times 32$, or $64 \times 64$).
- **Total Spots**: The total number of spots must be $L^2$ (a power of 4, such as 256, 1024, or 4096).
- **Coordinates (`coords`)**: A 2D array of shape `(N_spots, 2)` with integer indices from $0$ to $L-1$. Each grid location must appear exactly once.

### B. Expression Matrix (`X`)
- A 2D array of shape `(N_spots, N_genes)`.

---

## 2. Key Parameters

- `n_factors`: Number of latent spatial patterns to find (e.g. $5$ or $10$).
- `n_length_scales`: Number of wavelet detail levels ($D$). Must satisfy $2^D \le L$.
- `n_init`: Number of random initializations (WaveFactor keeps the best run based on ELBO).

---

## 3. What WaveFactor Returns (`result`)

After calling `model.fit(X, coords)`:

| Output | Shape | Meaning |
| :--- | :--- | :--- |
| `result.factors` | `(N_spots, K)` | Spatial factor values at each spot. |
| `result.spatial_factor_maps` | `(L, L, K)` | Reconstructed 2D spatial maps of each factor. |
| `result.loadings` | `(K, N_genes)` | How strongly each gene belongs to each factor. |
| `result.gene_pip` | `(K, N_genes)` | Gene inclusion probability ($\approx 1$ = active gene, $\approx 0$ = noise). |
| `result.spatial_pip` | List of arrays | Wavelet inclusion probabilities across spatial length scales. |
| `result.elbo` | `float` | Final model fit score (higher is better). |


---

# Installation Guide

> **Note**: This documentation is a **work in progress** and subject to further improvements.

---

## 1. Requirements

Before installing, make sure you have:
- **Python** $\ge 3.7$
- **CMake** $\ge 3.16$
- **A C++17 compiler** (e.g. GCC or Clang)

---

## 2. Install in an Isolated Environment

Using an isolated virtual environment prevents conflicts with other Python packages:

```bash
# 1. Clone this repository (copy URL from the green "Code" button on GitHub)
git clone https://github.com/shimlab/WaveFactor.git
cd WaveFactor

# 2. Create and activate a virtual environment
python3 -m venv wavefactor-venv
source wavefactor-venv/bin/activate

# 3. Install WaveFactor
pip install .
```

> **Warning** if installing into multiple environments or reinstalling, we recommend deleting the `build` folder created after every installation before starting a new installation so as to avoid potential conflicts with installation artifacts.

---

## 3. (Optional) Check Installation Succeeded

Check that the installation succeeded:

```bash
pip show wavefactor
```

---

## 4. (Optional) Deactivate or Remove

Deactivate the virtual environment by:

```bash
deactivate
```

After deactivation, to completely delete the installation, just remove the `wavefactor-venv/` folder:

```bash
rm -rf wavefactor-venv
```


---

