GEMS / Training Grounds
Four independent AI research lineages—generalist, software engineering, quantitative reasoning, and multimodal—plus the Training Grounds used to teach, evaluate, and advance them.
Current Research
- Topaz: broad language, reasoning, planning, and orchestration
- Sapphire: coding, repair, repository reasoning, and engineering tools
- Peridot: mathematics, science, formal reasoning, and verification
- Garnet: document and visual understanding, with image generation kept as a separate module
- Curriculum design, held-out evaluation, and learning-versus-memorization checks
- Hardware-aware training and evaluation
Long-Term Direction
As datasets, compute, funding, and research capacity expand, GEMS continues toward deeper FDS-developed, from-scratch foundations. Those are research directions—not shipped capability claims.
The Model Family
Meet the GEMS and Training Grounds
Where the program stands
The short version before the research detail: what generation the program is in, what is actually happening now, and what does not exist yet.
- Program stage
- Research program
- Active focus
- From-scratch model research — Training Grounds active
- Public model
- Not released
- Training status
- Research in progress
- Primary lineages
- Topaz · Sapphire · Peridot · Garnet
The GEMS family
GEMS is not one model wearing different hats. It is a family of independent research lineages: Topaz is the broad generalist and orchestrator, Sapphire focuses on software engineering, Peridot on mathematics and technical reasoning, and Garnet on multimodal and publishing intelligence.
GEM // TOPAZ
Topaz
General intelligence and orchestration
STATE // RESEARCH
Broad language, reasoning, planning, instruction following, and coordination across specialist systems.
Not claimed yetNo released Topaz model, frontier parity, or validated capability is claimed.
More about Topaz
GEM // SAPPHIRE
Sapphire
Software engineering and coding
STATE // RESEARCH
Code generation, repair, repository reasoning, navigation, testing, and structured engineering tool use.
Not claimed yetNo trained Sapphire model exists, and no Sapphire capability is presented as shipping in CodeForge.
More about Sapphire
GEM // PERIDOT
Peridot
Mathematics and technical reasoning
STATE // RESEARCH
Mathematics, formal and quantitative reasoning, science, structured problem solving, and verifiable technical work.
Not claimed yetNo trained Peridot model, independently verified benchmark result, or production capability is claimed.
More about Peridot
GEM // GARNET
Garnet
Multimodal and publishing intelligence
STATE // RESEARCH
Document and visual understanding, publishing workflows, and multimodal production with separately evaluated components.
Not claimed yetNo trained Garnet model exists, and image generation is not claimed as a current Garnet capability.
More about Garnet
How a GEM advances
Development is treated as one loop, not a launch. Each rung is proven with evidence before the next is attempted.
Set the Curriculum
Define the curriculum, tasks, and evaluation plan suited to the target role and its research stage.
Teach
Train the GEM on carefully prepared material and tasks appropriate to its developing role.
Test
Evaluate whether the model actually learned the intended skill instead of memorizing patterns or succeeding by accident.
Diagnose
Study failures, weak generalization, repetition, reasoning mistakes, context limits, and other measurable problems.
Refine
Adjust curriculum, post-training, evaluation, or model strategy based on the findings.
Expand
Increase capability, context, tool use, specialization, and real-world usefulness as each rung is proven.
Apply
Integrate what is mature enough into FDS systems and applications while research continues in Training Grounds.
Training Grounds
Training Grounds is where the family learns and is measured. It is a governed space for experimentation: GEMS models encounter controlled challenges, are evaluated against prepared tasks, fail safely, and are measured before anything advances. Advancement follows evidence, not activity — a GEM moves to the next rung only when results support it. Public descriptions stay high-level on purpose; the underlying research, datasets, and methods are not exposed.
- Controlled trials
- Held-out evaluation
- Failure analysis
- Checkpoint gating
- Model lineage boundaries
Capability direction
This is research, not a shipped product catalog. The lists below distinguish what is being developed from where the work is headed. Nothing here should be read as a claim that these capabilities are available today.
IN DEVELOPMENT
Actively being researched and built, not yet shipped as finished standalone capabilities.
- ReasoningMulti-step problem solving and general inference across the family.
- Research assistanceHelping people explore, summarize, and reason over material.
- Tool use & agentic executionCarrying out steps under human control, not autonomous decisions.
- Coding assistanceSupport for building and testing software.
- Writing & creative workDrafting, editing, and illustration support tied to publishing projects.
- Publishing productionMoving ideas toward finished, production-ready work.
- Knowledge retrievalGrounded answers from provided sources.
- Document & visual understandingGarnet's from-scratch research across text, images, video, and documents; capability is not yet validated.
LONG-TERM DIRECTION
Where the work is headed as each proven rung enables the next.
- Long-context understandingWorking across longer documents and sessions as models develop.
- Game & world developmentAssistance for interactive experiences, explored with KyraBlox.
- Image generationA separate, longer-term Garnet module direction, with no committed approach yet.
- Structured professional tasksHuman-supervised tooling, including medical-coding research framed as task support — not diagnosis or medical advice.
Affordable by design
One reason FDS is developing the GEMS family is economic access.
The destination is not "cheap AI." It is capable AI that ordinary people, creators, developers, families, small organizations, and communities can actually afford to use. We are working toward the depth and usefulness people associate with premium frontier assistants — deep reasoning, long-context work, capable creation, coding, and tool use — while researching how much of that experience can be delivered through smaller, specialized, efficient systems instead of permanently attaching every useful task to frontier-model pricing.
This is an aspirational experience target, not a claim of current parity with Claude Opus, GPT/Sol-class systems, or any other frontier model.
Latest Updates
All notes →GEMS can specialize before frontier-scale pretraining
Historical record: Phase 158 Generation 0 research into starting GEMS from open foundations. This strategy has since been superseded — see the editorial note below.
On GEMS: training and evaluation as connected work
Why the GEMS research direction treats a training run and the evidence it produces as one piece of work, not two.