DBSA-27B, an 8.3M-node scholarly graph, and long-context reasoning infrastructure
Frontier models are bottlenecked by training data — especially long-context. Models advertise huge context windows, but the data that teaches genuine long-range reasoning is scarce, expensive, and hard to validate at scale.
- —DBSA-27B — continued pretraining and instruction tuning of a 27B reasoning model on long-context reasoning-over-code corpora
- —Full-text scholarly knowledge graph on the OpenAlex spine: ~8.3M full-text nodes and 530.7M edges from arXiv, PMC, bioRxiv, and medRxiv
- —Hybrid dense/sparse retrieval and reranking for graph-aware RAG over real scientific text
- —A deep-research engine that synthesizes multi-source evidence with per-source quality scoring and citations
“The next capability jump in AI is not more parameters — it is better long-context data.”