/
User Name
user@example.com

The architecture landscape, in one place.

Two things get called “ML architecture”: the classical statistical model families, and the neural network wirings. Both are here, tagged by what they are and what data they eat.

Start from your problem

Architecture Decision Wizard

Answer three questions about your dataset, constraints, and problem to get instant tailored architecture recommendations.

Step 1: What is your primary data modality?

Select the input format and problem structure:

Side-by-Side Model Comparison

Compare mechanics, trade-offs, computation profiles, and caveats side-by-side.

PRESETS:

The ML Architecture Timeline (1805–2026)

Track the evolutionary leaps across statistical learning, deep networks, attention mechanisms, and state-space foundation models.

Study & Knowledge Checks

Master architecture trade-offs, mechanisms, and caveats for research, engineering, and ML system design interviews.

How to read any of this

“Architecture” is almost never a single brand name — it's a stack of independent choices. Two models with different names usually differ in exactly one row of this grid, which is why every 2025–26 LLM release looks like a small edit inside one transformer block.

The choice grid

AxisOptions you'll see
Token mixerconvolution, recurrence, attention, state space, long convolution, or a mix
Residual streamplain, residual, dense, hyper-connections / mHC
NormalizationBatchNorm, LayerNorm, RMSNorm, QK-norm, GroupNorm, DeepNorm
ActivationReLU, GELU, SiLU, GLU family, SwiGLU/GEGLU
Positional schemesinusoidal, learned, relative, RoPE (+ scaling), ALiBi, NoPE
Attention shapeMHA → MQA → GQA → MLA, sliding window, sparse, compressed
Sparsitydense, mixture-of-experts (top-k, fine-grained, shared experts)
Objectivesupervised, self-supervised, contrastive, masked, autoregressive, diffusion, RL

How to pick, in practice

  • Tabular data: gradient-boosted trees win most of the time. Reach for neural tabular models only when you have many columns, mixed modality, or lots of data.
  • Images: a modern CNN backbone or a ViT, pretrained. Detect/segment with the task-specific heads.
  • Text: a pretrained transformer you fine-tune or prompt. Don't train from scratch.
  • Long sequences / streaming at scale: the state-space and hybrid family exists exactly for this.
  • Few hundred rows: linear models, k-NN and trees. A deep net will overfit and you'll blame the data.
  • Anything generative: pick the family by output type — diffusion for images/audio, autoregressive for text/code, flows and VAEs when you need a likelihood or a smooth latent space.

Where to keep reading

About the links on each card

Architecture names are stable; paper URLs are not. Every entry stores its canonical paper title, and each title is resolved against public bibliographic indexes — preferring an arXiv landing page, then a DOI, then the publisher's page. Titles that cannot be matched confidently link to a targeted search instead, rather than sending you to a guessed identifier that may have rotted. The verifier also fetches a sample of the resolved URLs on every build to confirm they still resolve.

0 shortlisted