Define the objective and evidence
Establish the intended work, success conditions, unacceptable failures, and the role of qualified human judgment before selecting a model or method.
We define your goals, prepare materials and select a base model. We choose and evaluate adaptation methods, then prepare the results for operation. Explore the methodology and the glossary below.
The POLARIS approach starts with your goals and materials. We then select techniques, train, validate and prepare the model for release. Feedback informs each stage, so the method can adapt to your project.
Establish the intended work, success conditions, unacceptable failures, and the role of qualified human judgment before selecting a model or method.
Audit provenance, rights, sensitivity, language and domain coverage, then prepare the corpus, data splits, and text representation.
Compare candidate models against the material, language, domain, task, license, architecture, context, and operating constraints.
Use the least invasive combination that can meet the objective, distinguishing runtime knowledge from changes to model representation or behavior.
Model modification
Model modification / Post-training
Turn the selected model, material, and method into a versioned and reproducible training recipe.
Match the model, sequence length, optimizer state, batch, method, and experiment scale to an appropriate compute arrangement.
Compare the modified model with its baseline on unseen language, domain, task, risk, and retention evaluations before promotion.
Package the approved model, adapter, documentation, evaluation evidence, and serving requirements for the transition into operation.
Terms are listed alphabetically. Search by name, abbreviation, or wording from a definition.
Showing 59 of 59 terms
No terms match this search.
Measurable conditions a model or system must satisfy before it can be approved for the next phase.
A relatively small set of trained parameters attached to a base model to change its behavior without replacing all original weights.
The particular upstream model and revision selected as the starting point for a specialization project.
A defined collection of tasks, examples, metrics, and procedures used to compare models or versions.
The degree to which a model's confidence corresponds to its actual likelihood of being correct.
Loss of previously learned capabilities when a model is trained on new material.
A saved state from a training run, containing model or adapter weights and sometimes optimizer state for resuming training.
A token generated by the model as output; in supervised training, it may be a token on which training loss is applied.
The maximum amount of tokenized input and output a model can handle in one request.
An additional fine-tuning phase that starts from an already modified checkpoint rather than the original base model.
Additional language-model training on a substantial corpus to deepen representation of a language, domain, style, or knowledge distribution.
A structured body of text, documents, transcripts, images, or other material assembled for training, retrieval, or evaluation.
A method using multiple graphics processing units (GPUs), in which each device holds a model copy and processes a different part of the training batch.
The separation of examples into training, validation, and test sets with distinct roles in model development.
An architecture in which substantially all relevant parameters participate in processing each token.
Training a student model to reproduce useful behavior or probability patterns from a larger or otherwise different teacher model.
A common data-parallel implementation that synchronizes gradients among model replicas running on multiple graphics processing units (GPUs).
A numerical representation that places semantically related inputs near one another in a vector space.
The process of measuring a model, component, or complete system against defined examples and criteria.
An umbrella term for additional training that modifies a pretrained model for a behavior, task, language, domain, or preference.
A broadly trained model intended to serve as the basis for many downstream tasks.
Fine-tuning that can update all or most original model parameters, providing more capacity at greater compute and governance cost.
A calculated signal indicating how model parameters should change to reduce the selected loss.
The extent to which an answer is supported by the source material made available to the system.
Output that invents, misstates, or presents unsupported information as if it were reliable.
Examples deliberately excluded from training so they can test whether performance generalizes to unseen material.
A workflow in which qualified people create, review, correct, approve, or route data and model outputs.
A training setting selected by the team rather than learned directly by the model.
Running a trained model to produce outputs. Training changes parameters; inference uses the resulting parameters.
(large language model as a judge)
Benchmark and validate the result / Evaluation methodUsing one language model to compare, score, or classify outputs from another model or system.
A hyperparameter controlling the size of parameter updates during training.
The mathematical objective minimized or optimized during training.
(low-rank adaptation)
Choose the specialization approach / Model modification / Post-training / Parameter-efficient fine-tuningA parameter-efficient method that trains small low-rank components associated with selected model weights while leaving the original weights unchanged.
Two examples differing in one controlled feature so an evaluation can isolate a specific linguistic capability.
An architecture containing multiple expert components while activating only a subset for a given token.
A versioned deliverable produced by training or release preparation, such as an adapter, checkpoint, merged model, or quantized model.
Documentation describing a model version, intended use, training basis, evaluation, limitations, risks, license, and operating requirements.
Methods that divide one model across multiple graphics processing units (GPUs) because the complete model or training state does not fit on one device.
Learned numerical parameters that determine how a model transforms inputs into outputs.
A model whose trained weights are available under a stated license; this does not necessarily make all code, data, or uses open.
When a model improves on training examples but fails to generalize to unseen representative examples.
(parameter-efficient fine-tuning)
Choose the specialization approach / Model modification / Post-trainingA family of methods that updates or adds a relatively small number of parameters instead of retraining all model weights.
A form of model parallelism that assigns different groups of model layers to different devices.
Training or alignment performed after a model's broad pretraining phase.
(proximal policy optimization)
Choose the specialization approach / Model modification / Post-training / Reinforcement learningA reinforcement-learning algorithm that limits the size of policy updates while optimizing a reward signal.
Representing model weights or computations with lower numerical precision to reduce memory use and often improve inference efficiency.
An evaluation designed to detect whether a new model or system version damaged behavior that previously worked.
A training approach in which behavior is optimized using rewards derived from outcomes, preferences, rules, models, or people.
(retrieval-augmented generation)
Choose the specialization approach / Without weight modificationA system pattern that retrieves relevant external material at request time and supplies it to a generative model without modifying its weights.
A function, model, rule, or human judgment used to score behavior during reinforcement or preference training.
The number of tokens processed in one training example or request.
(supervised fine-tuning)
Choose the specialization approach / Model modification / Post-trainingTraining on labeled examples or demonstrations that show desired inputs and outputs.
Training or evaluation examples generated rather than directly collected from real events or documents.
A form of model parallelism that divides individual mathematical operations or tensors across multiple graphics processing units (GPUs).
A unit into which a tokenizer divides text for model processing; it may be a word, part of a word, punctuation, or another symbol.
The component that converts text into token identifiers and back, affecting efficiency for particular languages and scripts.
The reproducible configuration for a run: data, base model, method, loss, optimizer, hyperparameters, code, hardware, and evaluation schedule.
Held-out examples used during experimentation to compare recipes and tune development choices.
Distributed-training memory optimizations that partition optimizer state, gradients, and model parameters across devices.
Need a model specialized for your language and business?
With your permission, Google Analytics helps us understand site visits, and Microsoft Clarity shows interactions through heatmaps and session recordings. These tools use cookies. Neither loads until you allow it.
You’ll find our privacy notice and privacy settings in the footer.
Changing an active choice reloads this page. Save any inquiry text before changing it.
Your browser could not save this choice. Analytics stays off. You can continue using the site.