🛠️ Ambiakshi SLM Forge Studio

Research-grade Small Language Model (SLM) foundry for long-context financial sentiment, earnings transcript ingestion, and autonomous stock analysis.

1
Data Studio
PhraseBank & 8k Packing
2
Architecture & VRAM
Gemma 4 & 16GB P100 Matrix
3
Execution Hub
Live Telemetry & Loss
4
Evaluation & Benchmarks
F1, Confusion & Bias
5
Deployment & Hub
Hugging Face & Ollama

1. Dataset Ingestion & Stratification Dataset Verified

Select data preparation mode for Financial PhraseBank (all-data.csv) or custom financial disclosures.

1,363
Positive Sentences
2,879
Neutral Sentences
604
Negative Sentences
📁 Source: data/raw/all-data.csv (672 KB)
🔒 SHA-256 Checksum: 94a1b028... [Verified Match]
🌐 Encoding: ISO-8859-1 (Latin-1) -> Auto-normalized to UTF-8 Unix
🧩 Formatting: ChatML <start_of_turn> with Financial CoT JSON Output

💾 16GB P100 VRAM Profiler

Dynamic Backpropagation Memory Model
10.60 GB / 16.0 GB
Base Model: 4.20 GB
LoRA & Opt: 0.18 GB
8k Activations: 5.80 GB
CUDA Overhead: 1.15 GB
✅ 100% Kaggle P100 Compatible (Zero-OOM) +5.40 GB Headroom

🔬 Mathematical Memory Proof:

Using Gradient Checkpointing, activation memory drops from $O(N \cdot L) \approx 24\text{ GB}$ to $O(\sqrt{N}) \approx 5.8\text{ GB}$. Unquantized 16-bit base weights ($4.2\text{ GB}$) fit comfortably under the 16.0 GB hardware ceiling without 4-bit loss.