INTELLIGENCERESEARCH

GEMS / Training Grounds

Four independent AI research lineages—generalist, software engineering, quantitative reasoning, and multimodal—plus the Training Grounds used to teach, evaluate, and advance them.

From-Scratch ResearchCurriculumEvaluationAgentic LearningModel Lineages
Diagram showing Training Grounds as the shared learning and evaluation program at the center, connected to four independent GEMS research lineages: Topaz, Sapphire, Peridot, and Garnet.
Research artifactGEMS learning architecture: one shared Training Grounds program, four independent from-scratch research lineages.

Current Research

  • Topaz: broad language, reasoning, planning, and orchestration
  • Sapphire: coding, repair, repository reasoning, and engineering tools
  • Peridot: mathematics, science, formal reasoning, and verification
  • Garnet: document and visual understanding, with image generation kept as a separate module
  • Curriculum design, held-out evaluation, and learning-versus-memorization checks
  • Hardware-aware training and evaluation

Long-Term Direction

As datasets, compute, funding, and research capacity expand, GEMS continues toward deeper FDS-developed, from-scratch foundations. Those are research directions—not shipped capability claims.

The Model Family

Meet the GEMS and Training Grounds

Where the program stands

The short version before the research detail: what generation the program is in, what is actually happening now, and what does not exist yet.

Program stage
Research program
Active focus
From-scratch model research — Training Grounds active
Public model
Not released
Training status
Research in progress
Primary lineages
Topaz · Sapphire · Peridot · Garnet

The GEMS family

GEMS is not one model wearing different hats. It is a family of independent research lineages: Topaz is the broad generalist and orchestrator, Sapphire focuses on software engineering, Peridot on mathematics and technical reasoning, and Garnet on multimodal and publishing intelligence.

GEM // TOPAZ

Topaz

General intelligence and orchestration

STATE // RESEARCH

Broad language, reasoning, planning, instruction following, and coordination across specialist systems.

Not claimed yetNo released Topaz model, frontier parity, or validated capability is claimed.

More about Topaz
Current research status
No checkpoint has been trained or evaluated for Topaz. Curriculum and evaluation design for its generalist, orchestration-focused role is in early Training Grounds research.
May eventually help with
Become the broadest GEMS generalist and orchestrate specialists where focused expertise is more useful than one model doing everything.
Where it fits
The broad generalist in a family of independent lineages—not a shared base that every other GEM inherits.

GEM // SAPPHIRE

Sapphire

Software engineering and coding

STATE // RESEARCH

Code generation, repair, repository reasoning, navigation, testing, and structured engineering tool use.

Not claimed yetNo trained Sapphire model exists, and no Sapphire capability is presented as shipping in CodeForge.

More about Sapphire
Current research status
No checkpoint has been trained or evaluated for Sapphire. Its software-engineering curriculum, repository-reasoning tasks, and evaluation design are in early Training Grounds research.
May eventually help with
Develop into the software-engineering specialist intended to support CodeForge-oriented research and other repository work.
Where it fits
The coding specialist; its role is distinct from CodeForge, which is already a public engineering product.

GEM // PERIDOT

Peridot

Mathematics and technical reasoning

STATE // RESEARCH

Mathematics, formal and quantitative reasoning, science, structured problem solving, and verifiable technical work.

Not claimed yetNo trained Peridot model, independently verified benchmark result, or production capability is claimed.

More about Peridot
Current research status
No checkpoint has been trained or evaluated for Peridot. Its mathematics and technical-reasoning curriculum, with programmatic and formal verification, is in early Training Grounds research.
May eventually help with
Pursue correctness-first reasoning with programmatic, symbolic, and formal verification where appropriate.
Where it fits
The quantitative specialist. Training Grounds—not Peridot itself—owns the shared evaluation and advancement discipline.

GEM // GARNET

Garnet

Multimodal and publishing intelligence

STATE // RESEARCH

Document and visual understanding, publishing workflows, and multimodal production with separately evaluated components.

Not claimed yetNo trained Garnet model exists, and image generation is not claimed as a current Garnet capability.

More about Garnet
Current research status
No checkpoint has been trained or evaluated for Garnet. Its document, vision, and publishing-oriented curriculum is in early Training Grounds research; image generation remains a separate, longer-term module direction with no committed approach yet.
May eventually help with
Develop a multimodal system that can support document, vision, publishing, and image-generation workflows without conflating unlike model components.
Where it fits
The multimodal specialist, with potential relevance to Kayla Publisher while remaining a separate research lineage and system.

How a GEM advances

Development is treated as one loop, not a launch. Each rung is proven with evidence before the next is attempted.

  1. Set the Curriculum

    Define the curriculum, tasks, and evaluation plan suited to the target role and its research stage.

  2. Teach

    Train the GEM on carefully prepared material and tasks appropriate to its developing role.

  3. Test

    Evaluate whether the model actually learned the intended skill instead of memorizing patterns or succeeding by accident.

  4. Diagnose

    Study failures, weak generalization, repetition, reasoning mistakes, context limits, and other measurable problems.

  5. Refine

    Adjust curriculum, post-training, evaluation, or model strategy based on the findings.

  6. Expand

    Increase capability, context, tool use, specialization, and real-world usefulness as each rung is proven.

  7. Apply

    Integrate what is mature enough into FDS systems and applications while research continues in Training Grounds.

Training Grounds

Training Grounds is where the family learns and is measured. It is a governed space for experimentation: GEMS models encounter controlled challenges, are evaluated against prepared tasks, fail safely, and are measured before anything advances. Advancement follows evidence, not activity — a GEM moves to the next rung only when results support it. Public descriptions stay high-level on purpose; the underlying research, datasets, and methods are not exposed.

  • Controlled trials
  • Held-out evaluation
  • Failure analysis
  • Checkpoint gating
  • Model lineage boundaries
Training Grounds conceptual process diagram showing curriculum, training, checkpoint, evaluation, failure analysis, replication, capability gate, and advance or revise.
Conceptual diagramTraining Grounds research workflow schematic derived from the current project description.

Capability direction

This is research, not a shipped product catalog. The lists below distinguish what is being developed from where the work is headed. Nothing here should be read as a claim that these capabilities are available today.

IN DEVELOPMENT

Actively being researched and built, not yet shipped as finished standalone capabilities.

  • ReasoningMulti-step problem solving and general inference across the family.
  • Research assistanceHelping people explore, summarize, and reason over material.
  • Tool use & agentic executionCarrying out steps under human control, not autonomous decisions.
  • Coding assistanceSupport for building and testing software.
  • Writing & creative workDrafting, editing, and illustration support tied to publishing projects.
  • Publishing productionMoving ideas toward finished, production-ready work.
  • Knowledge retrievalGrounded answers from provided sources.
  • Document & visual understandingGarnet's from-scratch research across text, images, video, and documents; capability is not yet validated.

LONG-TERM DIRECTION

Where the work is headed as each proven rung enables the next.

  • Long-context understandingWorking across longer documents and sessions as models develop.
  • Game & world developmentAssistance for interactive experiences, explored with KyraBlox.
  • Image generationA separate, longer-term Garnet module direction, with no committed approach yet.
  • Structured professional tasksHuman-supervised tooling, including medical-coding research framed as task support — not diagnosis or medical advice.

Affordable by design

One reason FDS is developing the GEMS family is economic access.

The destination is not "cheap AI." It is capable AI that ordinary people, creators, developers, families, small organizations, and communities can actually afford to use. We are working toward the depth and usefulness people associate with premium frontier assistants — deep reasoning, long-context work, capable creation, coding, and tool use — while researching how much of that experience can be delivered through smaller, specialized, efficient systems instead of permanently attaching every useful task to frontier-model pricing.

This is an aspirational experience target, not a claim of current parity with Claude Opus, GPT/Sol-class systems, or any other frontier model.

Latest Updates

All notes →