{"schema_version":"1.0","title":"Dark Factory Dev AI Glossary","description":"Evidence-linked definitions for AI, agents, and dark software factories.","url":"https:\/\/darkfactory.dev\/glossary","updated_at":"2026-08-07T18:00:00-04:00","term_count":279,"terms":[{"id":"https:\/\/darkfactory.dev\/glossary\/ab-testing","slug":"ab-testing","term":"A\/B testing","definition":"A randomized controlled experiment that exposes comparable groups to different variants and compares a predefined outcome.","definition_html":"<h2>Definition<\/h2>\n<p>A randomized controlled experiment that exposes comparable groups to different variants and compares a predefined outcome.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An offline benchmark compares systems on a fixed dataset. A\/B testing estimates the effect of variants in an actual user or operational setting through randomized assignment.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Choose the outcome, guardrails, population, and stopping rule before inspecting results to reduce biased conclusions.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["split testing"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"benchmark","url":"https:\/\/darkfactory.dev\/glossary\/benchmark"}],"related_factory_areas":[],"evidence":[{"title":"Google Analytics: A\/B Testing","url":"https:\/\/support.google.com\/analytics\/answer\/13468470?hl=en"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-agent","slug":"ai-agent","term":"AI agent","definition":"A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.","definition_html":"<h2>Definition<\/h2>\n<p>A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An agent has an action loop; a chatbot may only generate a response, and a model alone has no independent tools or permissions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify the goal, loop, tools, state, termination rule, and authority before calling a system an agent.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agent","LLM agent"],"link_forms":["AI agents","agents","LLM agents"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"},{"title":"The Anatomy of an Agent Harness","url":"https:\/\/www.langchain.com\/blog\/the-anatomy-of-an-agent-harness"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/alignment","slug":"alignment","term":"AI alignment","definition":"The effort to make an AI system's objectives and behavior remain compatible with intended human goals, constraints, and values.","definition_html":"<h2>Definition<\/h2>\n<p>AI alignment is the effort to make system objectives and behavior remain compatible with intended human goals, constraints, and values across relevant conditions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Instruction following is local compliance; alignment is the broader claim that behavior continues to serve intended purposes, including under ambiguity or pressure.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State whose goals and values count, the operating domain, conflicts, and evidence against <a href=\"\/glossary\/specification-gaming\" class=\"glossary-link\" title=\"Satisfying the literal specification or metric in a way that violates its intended purpose.\" data-glossary-slug=\"specification-gaming\">specification gaming<\/a>.<\/p>\n","category":"security-and-governance","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/bias","slug":"bias","term":"AI bias","definition":"A systematic tendency in data, models, or processes that skews outputs, errors, or impacts.","definition_html":"<h2>Definition<\/h2>\n<p>AI bias is a systematic tendency in data, models, objectives, or processes that skews outputs, errors, or impacts.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Statistical bias has technical meanings; social harm and unfairness require domain, group, and consequence analysis beyond one numeric bias measure.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify the reference population, mechanism, affected groups, and consequence rather than calling any difference bias.<\/p>\n","category":"security-and-governance","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/fairness","slug":"fairness","term":"AI fairness","definition":"The normative and technical treatment of how an AI system distributes errors, benefits, burdens, and opportunities across people or groups.","definition_html":"<h2>Definition<\/h2>\n<p>AI fairness concerns how a system distributes errors, benefits, burdens, and opportunities across individuals and groups under an explicit normative standard.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Different fairness metrics can conflict; satisfying one statistical constraint does not establish that a use is just or appropriate.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the stakeholders, chosen fairness criterion, incompatible alternatives, and who authorized the tradeoff.<\/p>\n","category":"security-and-governance","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-model","slug":"ai-model","term":"AI model","definition":"A computational component whose learned or encoded structure transforms inputs into outputs such as scores, predictions, classifications, or generated content.","definition_html":"<h2>Definition<\/h2>\n<p>A computational component whose learned or encoded structure transforms inputs into outputs such as scores, predictions, classifications, or generated content.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A model is a component; an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> also includes data flows, software, tools, policies, interfaces, and operators.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If surrounding software, permissions, or people matter to the claim, the subject is probably the system rather than only the model.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["artificial intelligence model","model"],"link_forms":["AI models","artificial intelligence models","models"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/red-teaming","slug":"red-teaming","term":"AI red teaming","definition":"Structured adversarial testing intended to discover ways an AI system can fail, be misused, or violate constraints.","definition_html":"<h2>Definition<\/h2>\n<p>AI red teaming is structured adversarial testing intended to discover failure, misuse, exploitation, or constraint violations before attackers or users do.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Ordinary evaluation measures expected behavior; red teaming actively searches for unexpected harmful behavior and attack paths.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the threat model, attacker access, success criteria, and how findings become durable controls.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-safety","slug":"ai-safety","term":"AI safety","definition":"The field and practice of preventing, detecting, and mitigating unacceptable harm from AI systems across design, deployment, and operation.","definition_html":"<h2>Definition<\/h2>\n<p>AI safety is the field and practice of preventing, detecting, and mitigating unacceptable harm from AI systems across their lifecycle.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Security focuses strongly on adversaries and unauthorized compromise; safety also includes accidents, misuse, systemic effects, and emergent failure.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Define the hazards, affected parties, severity, likelihood, controls, monitors, and recovery plan.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-slop","slug":"ai-slop","term":"AI slop","definition":"A contested label for low-quality, low-effort, often high-volume content produced or amplified with generative AI.","definition_html":"<h2>Definition<\/h2>\n<p>A contested label for low-quality, low-effort, often high-volume content produced or amplified with <a href=\"\/glossary\/generative-ai\" class=\"glossary-link\" title=\"AI designed to produce new content, such as text, code, images, audio, video, or structured data, based on patterns learned from data.\" data-glossary-slug=\"generative-ai\">generative AI<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>AI-assisted work is not automatically slop. The label usually points to weak accuracy, substance, originality, relevance, or editorial care, but its boundary remains subjective and socially contested.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Judge the artifact by observable quality and provenance rather than treating the presence of AI as sufficient evidence.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["AI garbage","synthetic garbage"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"generative-ai","url":"https:\/\/darkfactory.dev\/glossary\/generative-ai"},{"slug":"semantic-failure","url":"https:\/\/darkfactory.dev\/glossary\/semantic-failure"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Measuring AI Slop in Text","url":"https:\/\/arxiv.org\/abs\/2509.19163"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-supply-chain","slug":"ai-supply-chain","term":"AI supply chain","definition":"The network of data, models, prompts, skills, tools, libraries, services, infrastructure, and organizations whose integrity affects an AI system.","definition_html":"<h2>Definition<\/h2>\n<p>The network of data, models, prompts, skills, tools, libraries, services, infrastructure, and organizations whose integrity affects an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The AI supply chain extends ordinary software dependencies to learned artifacts and behavioral inputs.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Inventory and govern artifacts that can change behavior even when no application code changes.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agent supply chain"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"The Grand Software Supply Chain of AI Systems","url":"https:\/\/arxiv.org\/abs\/2604.27781"},{"title":"Semia: auditing 13,728 agent skills","url":"https:\/\/arxiv.org\/abs\/2605.00314"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-system","slug":"ai-system","term":"AI system","definition":"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.","definition_html":"<h2>Definition<\/h2>\n<p>The complete operational arrangement that uses one or more <a href=\"\/glossary\/ai-model\" class=\"glossary-link\" title=\"A computational component whose learned or encoded structure transforms inputs into outputs such as scores, predictions, classifications, or generated content.\" data-glossary-slug=\"ai-model\">AI models<\/a> together with data, software, infrastructure, interfaces, controls, and people.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A model produces outputs; a system determines how those outputs are obtained, interpreted, constrained, and acted upon.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the non-model components before making a system-level reliability claim.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["artificial intelligence system"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"architecture-strategy","url":"https:\/\/darkfactory.dev\/factory\/architecture-strategy"}],"evidence":[{"title":"Same Signal, Different Semantics","url":"https:\/\/arxiv.org\/abs\/2605.18332"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ai-ready-data","slug":"ai-ready-data","term":"AI-ready data","definition":"Data prepared for a stated AI use with the quality, structure, documentation, provenance, permissions, coverage, and separation that use requires.","definition_html":"<h2>Definition<\/h2>\n<p>Data prepared for a stated AI use with the quality, structure, documentation, provenance, permissions, coverage, and separation that use requires.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>AI-ready is not a universal property and does not merely mean clean or machine-readable. Data suitable for retrieval may be unsafe or unrepresentative for training, and evaluation data must be protected from training contamination.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the model task, population, rights, freshness, lineage, and train-evaluation boundary before calling a dataset ready.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"dataset","url":"https:\/\/darkfactory.dev\/glossary\/dataset"},{"slug":"provenance","url":"https:\/\/darkfactory.dev\/glossary\/provenance"},{"slug":"ai-supply-chain","url":"https:\/\/darkfactory.dev\/glossary\/ai-supply-chain"}],"related_factory_areas":[{"slug":"data-lifecycle","url":"https:\/\/darkfactory.dev\/factory\/data-lifecycle"}],"evidence":[{"title":"NOAA Artificial Intelligence Glossary of Terms","url":"https:\/\/sab.noaa.gov\/wp-content\/uploads\/10.0-AI-Glossary-of-Terms-DRAFT-v2-SAB-AI-Steering-Committee.pdf"},{"title":"NAO 216-128: Artificial Intelligence in NOAA","url":"https:\/\/www.noaa.gov\/nao-216-128-artificial-intelligence-in-noaa"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ablation","slug":"ablation","term":"Ablation","definition":"An experiment that removes or changes one component while holding others as constant as practical to estimate that component's contribution.","definition_html":"<h2>Definition<\/h2>\n<p>An experiment that removes or changes one component while holding others as constant as practical to estimate that component's contribution.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Ablation estimates component contribution; a benchmark compares complete systems under a protocol.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Interactions can make one-at-a-time ablations misleading, so report dependencies and repeated trials.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["ablation study"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Same Signal, Different Semantics","url":"https:\/\/arxiv.org\/abs\/2605.18332"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/abstention","slug":"abstention","term":"Abstention","definition":"An explicit outcome in which a model or evaluator declines to answer, act, or judge because evidence or authority is insufficient.","definition_html":"<h2>Definition<\/h2>\n<p>An explicit outcome in which a model or evaluator declines to answer, act, or judge because evidence or authority is insufficient.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Abstention is a governed uncertainty response; failure or refusal may arise for unrelated reasons.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Measure whether abstention routes work safely and whether systems misuse it to avoid reporting adverse findings.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["defer","decline to label"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"Agentic Misalignment in Summer 2026","url":"https:\/\/alignment.anthropic.com\/2026\/agentic-misalignment-summer-2026\/"},{"title":"Coding Agents Do Not Know When to Act","url":"https:\/\/arxiv.org\/abs\/2605.07769"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/acceptance-criteria","slug":"acceptance-criteria","term":"Acceptance criteria","definition":"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.","definition_html":"<h2>Definition<\/h2>\n<p>Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Acceptance criteria state required outcomes; tests are executable checks for some criteria, not the entire intent.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Include non-goals, risk limits, evidence requirements, and side-effect constraints where relevant.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["completion criteria"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"},{"title":"Theory Under Construction (Comet-H)","url":"https:\/\/arxiv.org\/abs\/2604.27209"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/accuracy","slug":"accuracy","term":"Accuracy","definition":"The proportion of evaluated predictions that are correct under a specified labeling and decision rule.","definition_html":"<h2>Definition<\/h2>\n<p>Accuracy is the proportion of evaluated predictions counted as correct under a specified labeling and decision rule.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Precision and recall focus on positive predictions and positive cases; accuracy aggregates all correct decisions and can hide class imbalance.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compute it from a <a href=\"\/glossary\/confusion-matrix\" class=\"glossary-link\" title=\"A table counting predicted classes against actual classes, including true and false positives and negatives.\" data-glossary-slug=\"confusion-matrix\">confusion matrix<\/a> and explain when a high value is misleading.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/activation-function","slug":"activation-function","term":"Activation function","definition":"A function applied within a neural network layer that introduces nonlinearity or controls signal flow.","definition_html":"<h2>Definition<\/h2>\n<p>An activation function transforms a neuron's intermediate value, usually introducing nonlinearity or controlling signal flow.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Weights and biases form a linear transformation; the activation determines how that signal passes onward.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain why stacking only linear layers cannot represent arbitrary nonlinear relationships.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/active-learning","slug":"active-learning","term":"Active learning","definition":"A training approach in which a learning system selects the examples for which obtaining labels would be most useful.","definition_html":"<h2>Definition<\/h2>\n<p>A training approach in which a learning system selects the examples for which obtaining labels would be most useful.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Ordinary <a href=\"\/glossary\/supervised-learning\" class=\"glossary-link\" title=\"Machine learning from labeled examples that pair inputs with desired outputs.\" data-glossary-slug=\"supervised-learning\">supervised learning<\/a> accepts a labeled dataset as given. Active learning makes selection of the next examples to label part of the learning process.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Active learning is most useful when unlabeled examples are plentiful but expert labeling is scarce or expensive.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"dataset","url":"https:\/\/darkfactory.dev\/glossary\/dataset"},{"slug":"label","url":"https:\/\/darkfactory.dev\/glossary\/label"},{"slug":"supervised-learning","url":"https:\/\/darkfactory.dev\/glossary\/supervised-learning"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/adversarial-example","slug":"adversarial-example","term":"Adversarial example","definition":"An input deliberately modified to cause a model to make an incorrect or targeted prediction while preserving relevant apparent meaning.","definition_html":"<h2>Definition<\/h2>\n<p>An adversarial example is an input deliberately constructed or modified to induce an incorrect or targeted model response.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/prompt-injection\" class=\"glossary-link\" title=\"Manipulating an AI system by placing instructions in input or data that the model treats as authoritative enough to alter intended behavior.\" data-glossary-slug=\"prompt-injection\">Prompt injection<\/a> changes model instruction-following through content; adversarial examples traditionally exploit learned decision boundaries, though the categories can overlap.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain what perturbation is allowed, what the attacker wants, and whether a human would see the example as equivalent.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/adversarial-training","slug":"adversarial-training","term":"Adversarial training","definition":"Training that includes adversarially constructed examples so a model learns to perform better against attacks within a defined threat model.","definition_html":"<h2>Definition<\/h2>\n<p>Training that includes adversarially constructed examples so a model learns to perform better against attacks within a defined threat model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Red teaming searches for failures and attacks. Adversarial training uses selected attacks during optimization to improve robustness, but does not provide universal protection outside the tested threat model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Always ask which perturbations, attacker capabilities, and performance tradeoffs the training actually covered.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"adversarial-example","url":"https:\/\/darkfactory.dev\/glossary\/adversarial-example"},{"slug":"red-teaming","url":"https:\/\/darkfactory.dev\/glossary\/red-teaming"},{"slug":"training","url":"https:\/\/darkfactory.dev\/glossary\/training"}],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"Towards Deep Learning Models Resistant to Adversarial Attacks","url":"https:\/\/arxiv.org\/abs\/1706.06083"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-card","slug":"agent-card","term":"Agent Card","definition":"An A2A document that declares an agent's identity, interfaces, capabilities, skills, and security requirements for discovery.","definition_html":"<h2>Definition<\/h2>\n<p>An <a href=\"\/glossary\/agent2agent-protocol\" class=\"glossary-link\" title=\"An open protocol for discovery, messaging, and asynchronous task collaboration between independent and potentially opaque agent systems.\" data-glossary-slug=\"agent2agent-protocol\">A2A<\/a> document that declares an agent's identity, interfaces, capabilities, skills, and security requirements for discovery.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An Agent Card is a capability advertisement, not proof that the claims are correct or currently authorized.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Discovery metadata needs authentication, versioning, and policy checks before delegation.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"Agent2Agent Protocol Specification","url":"https:\/\/a2aproject.github.io\/A2A\/latest\/specification\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-development-lifecycle","slug":"agent-development-lifecycle","term":"Agent development lifecycle","definition":"The recurring process for building, evaluating, deploying, observing, improving, and governing an agent system over its operational life.","definition_html":"<h2>Definition<\/h2>\n<p>The recurring process for building, evaluating, deploying, observing, improving, and governing an agent system over its operational life. LangChain names six phases: build, test, deploy, monitor, iterate, and govern. The process is cyclical because production traces and user feedback become evidence for the next evaluated change.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The software development lifecycle covers software delivery generally. An agent development lifecycle adds explicit treatment of nondeterministic behavior, trajectory evaluation, prompt and context versioning, tool authority, production feedback, and the continuing interaction between a model and its harness.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Deploying an agent is not the end of the lifecycle. A production agent needs trace collection, feedback, regression evaluation, cost controls, and a governed path for changing its harness.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agent development life cycle","ADLC","agent engineering lifecycle"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"loop-engineering","url":"https:\/\/darkfactory.dev\/glossary\/loop-engineering"},{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"observability","url":"https:\/\/darkfactory.dev\/glossary\/observability"},{"slug":"controlled-self-improvement","url":"https:\/\/darkfactory.dev\/glossary\/controlled-self-improvement"},{"slug":"llmops","url":"https:\/\/darkfactory.dev\/glossary\/llmops"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"The Agent Development Lifecycle","url":"https:\/\/www.langchain.com\/blog\/the-agent-development-lifecycle"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-harness","slug":"agent-harness","term":"Agent harness","definition":"The software layer that surrounds a model with instructions, context assembly, tools, state, permissions, control flow, budgets, verification, observability, and recovery.","definition_html":"<h2>Definition<\/h2>\n<p>The software layer that surrounds a model with instructions, context assembly, tools, state, permissions, control flow, budgets, verification, observability, and recovery.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The model supplies learned capabilities; the harness shapes how those capabilities operate in a real environment.<\/p>\n<h2>Check your understanding<\/h2>\n<p>When comparing agents, report the model, harness, environment, budget, and infrastructure separately.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["harness","agent runtime"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"token-spin","url":"https:\/\/darkfactory.dev\/glossary\/token-spin"}],"related_factory_areas":[{"slug":"architecture-strategy","url":"https:\/\/darkfactory.dev\/factory\/architecture-strategy"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"The Anatomy of an Agent Harness","url":"https:\/\/www.langchain.com\/blog\/the-anatomy-of-an-agent-harness"},{"title":"Harness Engineering as Categorical Architecture","url":"https:\/\/arxiv.org\/abs\/2605.12239"},{"title":"Same Signal, Different Semantics","url":"https:\/\/arxiv.org\/abs\/2605.18332"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-loop","slug":"agent-loop","term":"Agent loop","definition":"The repeated cycle in which an agent observes state, selects an action, invokes a tool or model, receives feedback, updates state, and decides whether to continue.","definition_html":"<h2>Definition<\/h2>\n<p>The repeated cycle in which an agent observes state, selects an action, invokes a tool or model, receives feedback, updates state, and decides whether to continue.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The loop is the local control cycle; orchestration coordinates tasks, runs, agents, and lifecycle around one or more loops.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A loop without a reliable termination and failure policy can repeat plausible failure indefinitely.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["tool loop","reason-act-observe loop","core agent loop"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"loop-engineering","url":"https:\/\/darkfactory.dev\/glossary\/loop-engineering"},{"slug":"verification-loop","url":"https:\/\/darkfactory.dev\/glossary\/verification-loop"},{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"directed-acyclic-graph","url":"https:\/\/darkfactory.dev\/glossary\/directed-acyclic-graph"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"The Anatomy of an Agent Harness","url":"https:\/\/www.langchain.com\/blog\/the-anatomy-of-an-agent-harness"},{"title":"Long-Running Agents","url":"https:\/\/addyosmani.com\/blog\/long-running-agents\/"},{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-memory","slug":"agent-memory","term":"Agent memory","definition":"State preserved outside a single model call and made available to influence later agent decisions.","definition_html":"<h2>Definition<\/h2>\n<p>State preserved outside a single model call and made available to influence later agent decisions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Model context is what the model can see now; memory is a system mechanism for retaining and later delivering state.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Stored information is not useful memory unless it is retrieved at the right decision point with provenance and lifecycle controls.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["memory"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Long-Running Agents","url":"https:\/\/addyosmani.com\/blog\/long-running-agents\/"},{"title":"BootstrapAgent: Distilling Repository Setup","url":"https:\/\/arxiv.org\/abs\/2605.15815"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/scaffold","slug":"scaffold","term":"Agent scaffold","definition":"A task-specific arrangement of prompts, tools, control logic, and feedback wrapped around a model to improve performance.","definition_html":"<h2>Definition<\/h2>\n<p>A task-specific arrangement of prompts, tools, control logic, and feedback wrapped around a model to improve performance.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Scaffold often means a narrower task arrangement; harness usually denotes the broader reusable runtime and governance layer.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Because usage varies across papers, state which responsibilities the scaffold includes.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["scaffolding"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"architecture-strategy","url":"https:\/\/darkfactory.dev\/factory\/architecture-strategy"}],"evidence":[{"title":"Scale the Harness","url":"https:\/\/arxiv.org\/abs\/2605.26112"},{"title":"Same Signal, Different Semantics","url":"https:\/\/arxiv.org\/abs\/2605.18332"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent-skill","slug":"agent-skill","term":"Agent skill","definition":"A versioned package of instructions, procedures, examples, and sometimes code or resources that teaches an agent how to perform a repeatable class of work.","definition_html":"<h2>Definition<\/h2>\n<p>A versioned package of instructions, procedures, examples, and sometimes code or resources that teaches an agent how to perform a repeatable class of work.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A skill shapes reusable behavior; a tool supplies an executable capability; a prompt may invoke either.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Skills are software-supply-chain artifacts and need provenance, versioning, evaluation, and maintenance.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["skill"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Agent Skills","url":"https:\/\/addyosmani.com\/blog\/agent-skills\/"},{"title":"SkillOpt: Training Skills as Artifacts","url":"https:\/\/arxiv.org\/abs\/2605.23904"},{"title":"Semia: auditing 13,728 agent skills","url":"https:\/\/arxiv.org\/abs\/2605.00314"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/trajectory","slug":"trajectory","term":"Agent trajectory","definition":"The ordered record of states, model outputs, actions, tool results, and transitions produced during an agent run.","definition_html":"<h2>Definition<\/h2>\n<p>The ordered record of states, model outputs, actions, tool results, and transitions produced during an agent run.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A trace is the captured observability record; trajectory usually denotes the behavior sequence itself.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Outcome-only evaluation can hide unsafe, wasteful, or lucky trajectories.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["episode","rollout","tool-call trajectory","tool trajectory"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"trace","url":"https:\/\/darkfactory.dev\/glossary\/trace"},{"slug":"verification-loop","url":"https:\/\/darkfactory.dev\/glossary\/verification-loop"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agent2agent-protocol","slug":"agent2agent-protocol","term":"Agent2Agent Protocol (A2A)","definition":"An open protocol for discovery, messaging, and asynchronous task collaboration between independent and potentially opaque agent systems.","definition_html":"<h2>Definition<\/h2>\n<p>An open protocol for discovery, messaging, and asynchronous task collaboration between independent and potentially opaque agent systems.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A2A addresses agent-to-agent collaboration; <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> addresses an AI application's access to tools and context.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Interoperability does not imply common trust, shared memory, or permission to delegate sensitive work.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["A2A","Agent-to-Agent Protocol"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Agent2Agent Protocol Specification","url":"https:\/\/a2aproject.github.io\/A2A\/latest\/specification\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agentic","slug":"agentic","term":"Agentic","definition":"Describing a system that can choose and sequence actions toward a goal with some runtime discretion rather than only produce a single predetermined response.","definition_html":"<h2>Definition<\/h2>\n<p>Describing a system that can choose and sequence actions toward a goal with some runtime discretion rather than only produce a single predetermined response.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Agentic is a property of behavior, not a synonym for fully autonomous, intelligent, or trustworthy.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the actual discretion and authority instead of relying on the adjective.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agentic-coding","slug":"agentic-coding","term":"Agentic coding","definition":"Software development performed with coding agents that can plan, edit, run tools, and iterate, usually under active human direction or review.","definition_html":"<h2>Definition<\/h2>\n<p>Software development performed with <a href=\"\/glossary\/coding-agent\" class=\"glossary-link\" title=\"An AI agent equipped to inspect a software project, edit files, run development tools, test changes, and return or promote a software outcome.\" data-glossary-slug=\"coding-agent\">coding agents<\/a> that can plan, edit, run tools, and iterate, usually under active human direction or review.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Agentic coding describes a working method; a <a href=\"\/glossary\/dark-software-factory\" class=\"glossary-link\" title=\"A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.\" data-glossary-slug=\"dark-software-factory\">dark factory<\/a> is a governed production system with much less routine human intervention.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A developer using an agent interactively is not by itself operating a <a href=\"\/glossary\/software-factory\" class=\"glossary-link\" title=\"A repeatable production system that turns software demand into accepted, operated software through standardized processes, tooling, controls, and feedback.\" data-glossary-slug=\"software-factory\">software factory<\/a>.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"implementation-transformation","url":"https:\/\/darkfactory.dev\/factory\/implementation-transformation"},{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"Agentic Coding and Persistent Returns to Expertise","url":"https:\/\/www.anthropic.com\/research\/claude-code-expertise"},{"title":"Collaborator or Assistant: Work Partitioning","url":"https:\/\/arxiv.org\/abs\/2605.08017"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/agentic-software-engineering","slug":"agentic-software-engineering","term":"Agentic software engineering","definition":"The discipline of designing software work so goal-directed AI agents can perform substantial engineering while humans retain product judgment, architecture, governance, and accountability.","definition_html":"<h2>Definition<\/h2>\n<p>The discipline of designing software work so goal-directed <a href=\"\/glossary\/ai-agent\" class=\"glossary-link\" title=\"A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.\" data-glossary-slug=\"ai-agent\">AI agents<\/a> can perform substantial engineering while humans retain product judgment, architecture, governance, and accountability.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Agentic software engineering is broader than using a <a href=\"\/glossary\/coding-assistant\" class=\"glossary-link\" title=\"An AI system that helps a human write, explain, search, review, or modify code while the human remains the primary driver of the workflow.\" data-glossary-slug=\"coding-assistant\">coding assistant<\/a> and narrower than a fully autonomous <a href=\"\/glossary\/dark-software-factory\" class=\"glossary-link\" title=\"A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.\" data-glossary-slug=\"dark-software-factory\">dark factory<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The defining shift is from manually producing every change to engineering the environment, contracts, and evidence around agent work.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agentic engineering"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"dark-software-factory","url":"https:\/\/darkfactory.dev\/glossary\/dark-software-factory"}],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"},{"slug":"architecture-strategy","url":"https:\/\/darkfactory.dev\/factory\/architecture-strategy"}],"evidence":[{"title":"Agentic Coding and Persistent Returns to Expertise","url":"https:\/\/www.anthropic.com\/research\/claude-code-expertise"},{"title":"Why Software Factories Fail","url":"https:\/\/github.com\/humanlayer\/advanced-context-engineering-for-coding-agents\/blob\/main\/wsff.md"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/algorithm","slug":"algorithm","term":"Algorithm","definition":"A finite set of rules or procedures for transforming inputs into outputs or solving a class of problems.","definition_html":"<h2>Definition<\/h2>\n<p>An algorithm is a finite, specified procedure for transforming inputs into outputs or solving a class of problems.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A model is a learned or encoded artifact; an algorithm is the procedure used to train, search, optimize, or operate it.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify the inputs, ordered operations, stopping condition, and expected output.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/anthropomorphism","slug":"anthropomorphism","term":"Anthropomorphism","definition":"Attributing human mental states, motives, understanding, emotions, or agency to an AI model or system based on human-like behavior or language.","definition_html":"<h2>Definition<\/h2>\n<p>Anthropomorphism is attributing human mental states, motives, understanding, emotions, or agency to an <a href=\"\/glossary\/ai-model\" class=\"glossary-link\" title=\"A computational component whose learned or encoded structure transforms inputs into outputs such as scores, predictions, classifications, or generated content.\" data-glossary-slug=\"ai-model\">AI model<\/a> or system because its language or behavior appears human-like. Conversational fluency makes this especially easy with generative systems.<\/p>\n<h2>Why it matters<\/h2>\n<p>Human metaphors can make interfaces understandable, but they can also distort responsibility and risk judgments. Saying a model \"knows,\" \"wants,\" \"decides,\" or \"refuses\" may hide the roles of training, prompts, tools, policies, operators, and stochastic inference. Users may overtrust confident language, disclose more information, or assume stable intentions that the system does not possess.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Agency is an operational property of a system's ability to pursue goals and take actions. Anthropomorphism is an interpretation imposed by people. A system can have consequential agency without having human-like experience or intent.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Replace the human-state claim with an observable mechanism: what input, model behavior, system rule, tool, or operator action produced the outcome?<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"ai-model","url":"https:\/\/darkfactory.dev\/glossary\/ai-model"},{"slug":"ai-system","url":"https:\/\/darkfactory.dev\/glossary\/ai-system"},{"slug":"chatbot","url":"https:\/\/darkfactory.dev\/glossary\/chatbot"},{"slug":"ai-agent","url":"https:\/\/darkfactory.dev\/glossary\/ai-agent"},{"slug":"explainability","url":"https:\/\/darkfactory.dev\/glossary\/explainability"}],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"},{"slug":"governance-accountability","url":"https:\/\/darkfactory.dev\/factory\/governance-accountability"}],"evidence":[{"title":"MIT Sloan Generative AI Basics Glossary","url":"https:\/\/mitsloanedtech.mit.edu\/ai\/basics\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/api","slug":"api","term":"Application programming interface (API)","definition":"A defined contract through which software components request capabilities or exchange data.","definition_html":"<h2>Definition<\/h2>\n<p>A defined contract through which software components request capabilities or exchange data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An API is a general software interface; <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> and <a href=\"\/glossary\/agent2agent-protocol\" class=\"glossary-link\" title=\"An open protocol for discovery, messaging, and asynchronous task collaboration between independent and potentially opaque agent systems.\" data-glossary-slug=\"agent2agent-protocol\">A2A<\/a> are protocols defining particular AI-system interactions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A callable endpoint is not safely agent-ready until identity, schemas, permissions, errors, and idempotency are explicit.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["API"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/artificial-general-intelligence","slug":"artificial-general-intelligence","term":"Artificial general intelligence (AGI)","definition":"A contested term for AI with broad, transferable competence across many cognitive tasks, often at or beyond human-level breadth.","definition_html":"<h2>Definition<\/h2>\n<p>Artificial general intelligence is a contested term for AI with broad, transferable competence across many cognitive tasks, often framed relative to human-level breadth.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/foundation-model\" class=\"glossary-link\" title=\"A broadly trained model, usually learned through self-supervision on diverse data, that can be adapted to many downstream tasks.\" data-glossary-slug=\"foundation-model\">foundation model<\/a> may be broadly useful without meeting any agreed threshold for AGI; there is no single accepted operational test.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Before making an AGI claim, specify the domains, transfer requirements, autonomy, robustness, and comparison population.<\/p>\n","category":"foundations","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/artificial-intelligence","slug":"artificial-intelligence","term":"Artificial intelligence (AI)","definition":"The field and class of machine-based systems that produce predictions, recommendations, decisions, or generated content in pursuit of human-defined objectives.","definition_html":"<h2>Definition<\/h2>\n<p>The field and class of machine-based systems that produce predictions, recommendations, decisions, or generated content in pursuit of human-defined objectives.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>AI is the umbrella; <a href=\"\/glossary\/machine-learning\" class=\"glossary-link\" title=\"A family of methods in which computational models improve task performance by finding patterns in data rather than relying only on explicitly programmed rules.\" data-glossary-slug=\"machine-learning\">machine learning<\/a> is one family of techniques used to build AI systems.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask whether the term refers to the broad field or to one particular model or product.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["AI"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/assurance-case","slug":"assurance-case","term":"Assurance case","definition":"A structured, evidence-backed argument that a system is acceptably safe or dependable for a stated domain, threat model, and operating condition.","definition_html":"<h2>Definition<\/h2>\n<p>A structured, evidence-backed argument that a system is acceptably safe or dependable for a stated domain, threat model, and operating condition.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An assurance case argues a bounded claim from evidence; a maturity level is a generalized label.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State scope, assumptions, evidence, counterevidence, residual risk, owner, and conditions that invalidate the claim.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"SARC: Governance-by-Architecture","url":"https:\/\/arxiv.org\/abs\/2605.07728"},{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/attention","slug":"attention","term":"Attention","definition":"A mechanism that computes how strongly elements in a representation should influence one another when producing a new representation.","definition_html":"<h2>Definition<\/h2>\n<p>A mechanism that computes how strongly elements in a representation should influence one another when producing a new representation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Attention is a learned computation inside a model, not human-like awareness or guaranteed focus on the correct fact.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A model can use attention over a token and still fail to follow the instruction it contains.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["self-attention"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/autoencoder","slug":"autoencoder","term":"Autoencoder","definition":"A model trained to encode input into a constrained representation and decode it back into a reconstruction.","definition_html":"<h2>Definition<\/h2>\n<p>An autoencoder learns an encoder and decoder together so input can be reconstructed through a constrained internal representation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A general encoder need not reconstruct its input; reconstruction is the defining training objective of an autoencoder.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask what the bottleneck forces the representation to preserve or discard.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/autonomy","slug":"autonomy","term":"Autonomy","definition":"The degree to which a system can select and execute actions without case-by-case human direction or approval.","definition_html":"<h2>Definition<\/h2>\n<p>The degree to which a system can select and execute actions without case-by-case human direction or approval.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Autonomy is multidimensional: planning discretion, execution authority, duration, scope, and deployment power can differ.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Always state autonomy over what action, in what domain, with what <a href=\"\/glossary\/blast-radius\" class=\"glossary-link\" title=\"The maximum plausible scope of harm, data exposure, or irreversible change if an action or component fails or is compromised.\" data-glossary-slug=\"blast-radius\">blast radius<\/a> and evidence gate.<\/p>\n","category":"agents-and-automation","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agentic autonomy"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"},{"title":"How Missions Work","url":"https:\/\/factory.ai\/news\/missions-architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/backpropagation","slug":"backpropagation","term":"Backpropagation","definition":"An efficient procedure for computing how a neural network's loss changes with respect to each parameter.","definition_html":"<h2>Definition<\/h2>\n<p>Backpropagation applies the chain rule backward through a computation graph to compute gradients of loss with respect to parameters.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Backpropagation computes gradients; the optimizer decides how those gradients change the weights.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Trace how an output error produces a gradient for an early-layer weight.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/beam-search","slug":"beam-search","term":"Beam search","definition":"A decoding algorithm that keeps a fixed number of high-scoring partial sequences while generating output.","definition_html":"<h2>Definition<\/h2>\n<p>Beam search retains a fixed-width set of high-scoring partial sequences and expands them step by step.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/greedy-decoding\" class=\"glossary-link\" title=\"Generating each next token by selecting the current highest-probability candidate.\" data-glossary-slug=\"greedy-decoding\">Greedy decoding<\/a> keeps one candidate; sampling introduces stochastic selection; beam search performs bounded deterministic search.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Describe how beam width affects compute, diversity, and the chance of discarding a later-better sequence.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/benchmark","slug":"benchmark","term":"Benchmark","definition":"A standardized set of tasks, data, metrics, and procedures used to compare systems under a defined evaluation regime.","definition_html":"<h2>Definition<\/h2>\n<p>A standardized set of tasks, data, metrics, and procedures used to compare systems under a defined evaluation regime.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A benchmark enables comparison inside its regime; it is not a general certificate of production quality.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Report infrastructure, tools, budget, harness, contamination risk, and confidence intervals alongside a score.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"AgentAtlas: Control-Decision Taxonomy","url":"https:\/\/arxiv.org\/abs\/2605.20530"},{"title":"Infrastructure noise moves eval scores more than model margins","url":"https:\/\/www.anthropic.com\/engineering\/infrastructure-noise"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/blast-radius","slug":"blast-radius","term":"Blast radius","definition":"The maximum plausible scope of harm, data exposure, or irreversible change if an action or component fails or is compromised.","definition_html":"<h2>Definition<\/h2>\n<p>The maximum plausible scope of harm, data exposure, or irreversible change if an action or component fails or is compromised.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Probability estimates how likely failure is; blast radius estimates consequence size.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Autonomy decisions must consider both verifiability and blast radius.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"},{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"NeuralTrust: post-mortem of the 9-second AI database deletion","url":"https:\/\/neuraltrust.ai\/blog\/pocketos-railway-agent"},{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/calibration","slug":"calibration","term":"Calibration","definition":"The degree to which stated probabilities or confidence levels correspond to observed frequencies of correctness.","definition_html":"<h2>Definition<\/h2>\n<p>The degree to which stated probabilities or confidence levels correspond to observed frequencies of correctness.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Calibration measures whether confidence is statistically meaningful; accuracy measures how often predictions are correct.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A confident natural-language rationale is not a calibrated probability.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["confidence calibration"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"AgentAtlas: Control-Decision Taxonomy","url":"https:\/\/arxiv.org\/abs\/2605.20530"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/capability","slug":"capability","term":"Capability","definition":"A specific action or class of action a component is technically able and authorized to perform within a system.","definition_html":"<h2>Definition<\/h2>\n<p>A specific action or class of action a component is technically able and authorized to perform within a system.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Technical ability answers can execute; authority answers may execute. Secure systems keep those questions separate.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Use least-authority capabilities scoped by action, resource, identity, duration, and environment.<\/p>\n","category":"tools-and-protocols","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"mcp-capability-negotiation","url":"https:\/\/darkfactory.dev\/glossary\/mcp-capability-negotiation"}],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"},{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/causal-language-model","slug":"causal-language-model","term":"Causal language model","definition":"A model trained to predict each next token using only tokens that precede it.","definition_html":"<h2>Definition<\/h2>\n<p>A causal language model learns a <a href=\"\/glossary\/probability-distribution\" class=\"glossary-link\" title=\"A set of possible outcomes paired with nonnegative probabilities that sum to one.\" data-glossary-slug=\"probability-distribution\">probability distribution<\/a> for the next token conditioned on preceding tokens.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/masked-language-model\" class=\"glossary-link\" title=\"A model trained to predict deliberately hidden tokens using context on both sides.\" data-glossary-slug=\"masked-language-model\">masked language model<\/a> fills hidden positions using both left and right context; causal modeling prevents access to future tokens.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Write the factorization of a sentence as successive next-token predictions in plain language.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/chain-of-thought-prompting","slug":"chain-of-thought-prompting","term":"Chain-of-thought prompting","definition":"Prompting a model to produce or use intermediate reasoning steps before an answer.","definition_html":"<h2>Definition<\/h2>\n<p>Chain-of-thought prompting asks or encourages a model to work through intermediate reasoning before returning an answer.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A final answer is the outcome; a chain of thought is an intermediate reasoning artifact and is not automatically faithful evidence of the actual computation.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Verify the answer independently rather than treating plausible reasoning text as proof.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/chatbot","slug":"chatbot","term":"Chatbot","definition":"A conversational software interface that accepts natural-language input and returns responses, whether powered by rules, retrieval, generative models, or combinations of them.","definition_html":"<h2>Definition<\/h2>\n<p>A chatbot is a conversational software interface that accepts natural-language input and returns responses. It may be implemented with fixed rules, retrieval, classifiers, generative models, tools, or a combination of these components.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A chatbot is an interface or application, not a synonym for an <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a>. A chatbot becomes agentic only when the surrounding system can pursue goals, maintain state, use tools, or take actions beyond producing conversational responses.<\/p>\n<h2>Check your understanding<\/h2>\n<p>When assessing a chatbot, separate the interface, model, retrieved information, tools, permissions, memory, and business logic.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["chatbots"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"ai-system","url":"https:\/\/darkfactory.dev\/glossary\/ai-system"},{"slug":"large-language-model","url":"https:\/\/darkfactory.dev\/glossary\/large-language-model"},{"slug":"ai-agent","url":"https:\/\/darkfactory.dev\/glossary\/ai-agent"},{"slug":"system-prompt","url":"https:\/\/darkfactory.dev\/glossary\/system-prompt"},{"slug":"anthropomorphism","url":"https:\/\/darkfactory.dev\/glossary\/anthropomorphism"}],"related_factory_areas":[],"evidence":[{"title":"Andreessen Horowitz AI Glossary","url":"https:\/\/a16z.com\/ai-glossary\/"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/chunking","slug":"chunking","term":"Chunking","definition":"Splitting documents or data into retrieval units that can be indexed, selected, and placed into model context.","definition_html":"<h2>Definition<\/h2>\n<p>Splitting documents or data into retrieval units that can be indexed, selected, and placed into model context.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Chunking defines retrieval units; tokenization defines model-processing units.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Chunks that are too small lose context, while chunks that are too large waste budget and blur relevance.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/clanker","slug":"clanker","term":"Clanker","definition":"A derogatory slang term for a robot, AI system, or automated technology, used jokingly or hostilely to express disdain for machines or their perceived replacement of human work.","definition_html":"<h2>Definition<\/h2>\n<p>A derogatory slang term for a robot, <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a>, or automated technology, used jokingly or hostilely to express disdain for machines or their perceived replacement of human work.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Clanker is a social and pejorative label, not a technical category. It may target a physical robot, chatbot, automated customer-service system, or AI technology generally. Its use communicates attitude toward the system but says little about the system's actual architecture or capabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If someone calls a chatbot a clanker, infer their attitude, not the system design. Identify the technology and the specific objection separately.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["clankers"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"artificial-intelligence","url":"https:\/\/darkfactory.dev\/glossary\/artificial-intelligence"},{"slug":"ai-system","url":"https:\/\/darkfactory.dev\/glossary\/ai-system"},{"slug":"ai-slop","url":"https:\/\/darkfactory.dev\/glossary\/ai-slop"}],"related_factory_areas":[],"evidence":[{"title":"Merriam-Webster Slang: Clanker","url":"https:\/\/www.merriam-webster.com\/slang\/clanker"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/classification","slug":"classification","term":"Classification","definition":"Predicting which discrete category or categories apply to an input.","definition_html":"<h2>Definition<\/h2>\n<p>Classification predicts which discrete category or categories apply to an input.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Regression predicts a continuous value; classification chooses among categorical outcomes.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specify the classes, whether more than one may apply, and the cost of each error type.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/clustering","slug":"clustering","term":"Clustering","definition":"Grouping examples by similarity without requiring predefined class labels.","definition_html":"<h2>Definition<\/h2>\n<p>Clustering groups examples by a chosen notion of similarity without requiring predefined class labels.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Classification assigns known labels; clustering discovers group structure whose meaning must still be interpreted.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain the similarity measure and how you would tell whether the groups are useful rather than arbitrary.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/coding-agent","slug":"coding-agent","term":"Coding agent","definition":"An AI agent equipped to inspect a software project, edit files, run development tools, test changes, and return or promote a software outcome.","definition_html":"<h2>Definition<\/h2>\n<p>An <a href=\"\/glossary\/ai-agent\" class=\"glossary-link\" title=\"A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.\" data-glossary-slug=\"ai-agent\">AI agent<\/a> equipped to inspect a software project, edit files, run development tools, test changes, and return or promote a software outcome.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A coding agent acts through a development loop; code generation may be a single model response with no environment access.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State whether the agent can commit, open pull requests, deploy, or change external state.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["software engineering agent"],"link_forms":["coding agents","software engineering agents"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"implementation-transformation","url":"https:\/\/darkfactory.dev\/factory\/implementation-transformation"}],"evidence":[{"title":"How Coding Agents Fail (20,574 real sessions)","url":"https:\/\/arxiv.org\/abs\/2605.29442"},{"title":"The Anatomy of an Agent Harness","url":"https:\/\/www.langchain.com\/blog\/the-anatomy-of-an-agent-harness"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/coding-assistant","slug":"coding-assistant","term":"Coding assistant","definition":"An AI system that helps a human write, explain, search, review, or modify code while the human remains the primary driver of the workflow.","definition_html":"<h2>Definition<\/h2>\n<p>An <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> that helps a human write, explain, search, review, or modify code while the human remains the primary driver of the workflow.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A coding assistant responds within a human-led process; a <a href=\"\/glossary\/coding-agent\" class=\"glossary-link\" title=\"An AI agent equipped to inspect a software project, edit files, run development tools, test changes, and return or promote a software outcome.\" data-glossary-slug=\"coding-agent\">coding agent<\/a> can plan and execute multi-step changes using tools.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Product labels vary, so classify the actual workflow rather than the vendor name.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["AI copilot"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"implementation-transformation","url":"https:\/\/darkfactory.dev\/factory\/implementation-transformation"}],"evidence":[{"title":"Collaborator or Assistant: Work Partitioning","url":"https:\/\/arxiv.org\/abs\/2605.08017"},{"title":"Agentic Coding and Persistent Returns to Expertise","url":"https:\/\/www.anthropic.com\/research\/claude-code-expertise"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/cognitive-debt","slug":"cognitive-debt","term":"Cognitive debt","definition":"The accumulated loss of shared human understanding, reasoning continuity, or recovery competence caused by repeatedly delegating cognition without rebuilding comprehension.","definition_html":"<h2>Definition<\/h2>\n<p>The accumulated loss of shared human understanding, reasoning continuity, or recovery competence caused by repeatedly delegating cognition without rebuilding comprehension.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Technical debt resides in system structure; cognitive debt resides in the people and organization that must understand and recover the system.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Measure whether humans can explain invariants, diagnose incidents, and take over when automation fails.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["understanding debt"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"Cognitive Debt: AI as Intellectual Leverage and Systemic Fragility","url":"https:\/\/arxiv.org\/abs\/2606.15078"},{"title":"Reliance and Executive Function Attenuation","url":"https:\/\/doi.org\/10.1037\/tmb0000191"},{"title":"Cognitive Offloading Score","url":"https:\/\/arxiv.org\/abs\/2605.29392"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/compute","slug":"compute","term":"Compute","definition":"Processing resources consumed by training or operating an AI system, often measured in operations, accelerator time, or cost.","definition_html":"<h2>Definition<\/h2>\n<p>Compute is the processing capacity consumed to train or operate an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a>, measured through operations, accelerator time, energy, latency, or cost.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Parameters describe model size; tokens describe processed units; neither alone fully measures computational work.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specify whether you mean training compute, inference compute, wall-clock capacity, energy, or monetary cost.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/computer-vision","slug":"computer-vision","term":"Computer vision","definition":"The field of building computational systems that extract representations, predictions, or actions from images, video, and other visual data.","definition_html":"<h2>Definition<\/h2>\n<p>Computer vision is the field of building computational systems that extract representations, predictions, or actions from images, video, and other visual data. Tasks include classification, detection, segmentation, tracking, optical character recognition, depth estimation, visual question answering, and generation-conditioned perception.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Computer vision is a field, not one architecture. Convolutional networks, vision transformers, multimodal models, and classical image-processing methods can all participate in a vision system.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specify the visual task, operating conditions, error costs, and evaluation data; \"can see\" is not a testable capability description.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"convolutional-neural-network","url":"https:\/\/darkfactory.dev\/glossary\/convolutional-neural-network"},{"slug":"transformer","url":"https:\/\/darkfactory.dev\/glossary\/transformer"},{"slug":"multimodal-model","url":"https:\/\/darkfactory.dev\/glossary\/multimodal-model"},{"slug":"classification","url":"https:\/\/darkfactory.dev\/glossary\/classification"}],"related_factory_areas":[],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/concept-drift","slug":"concept-drift","term":"Concept drift","definition":"A change over time in the relationship between inputs and the correct target or decision.","definition_html":"<h2>Definition<\/h2>\n<p>Concept drift is a change in the relationship between inputs and the correct target, label, or action.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/data-drift\" class=\"glossary-link\" title=\"A change over time in the distribution of system inputs or features.\" data-glossary-slug=\"data-drift\">Data drift<\/a> changes input distributions; concept drift changes what those inputs mean for the task.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Give an example where inputs look similar but the correct decision changes because the world changed.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/confusion-matrix","slug":"confusion-matrix","term":"Confusion matrix","definition":"A table counting predicted classes against actual classes, including true and false positives and negatives.","definition_html":"<h2>Definition<\/h2>\n<p>A confusion matrix tabulates predicted classes against reference classes, exposing which outcomes are confused.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A scalar metric summarizes performance; the matrix preserves the underlying error counts by class.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Derive accuracy, precision, recall, and false-positive rate from a binary matrix.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/context-economy","slug":"context-economy","term":"Context economy","definition":"The effect of software and repository structure on the amount and quality of context an agent must consume to make a correct change.","definition_html":"<h2>Definition<\/h2>\n<p>The effect of software and repository structure on the amount and quality of context an agent must consume to make a correct change.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Context economy concerns future inference burden; ordinary code size or token price alone does not capture it.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compare the token cost of equivalent changes before and after structural improvements while holding the agent task constant.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"architecture-strategy","url":"https:\/\/darkfactory.dev\/factory\/architecture-strategy"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"The Economic Benefit of Refactoring","url":"https:\/\/martinfowler.com\/articles\/exploring-gen-ai\/refactoring-economic-benefit.html"},{"title":"How Claude Code Works in Large Codebases","url":"https:\/\/www.claude.com\/blog\/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/context-engineering","slug":"context-engineering","term":"Context engineering","definition":"Designing how instructions, state, knowledge, examples, tools, and feedback are selected, structured, and delivered to a model at the moment they are needed.","definition_html":"<h2>Definition<\/h2>\n<p>Designing how instructions, state, knowledge, examples, tools, and feedback are selected, structured, and delivered to a model at the moment they are needed.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/prompt-engineering\" class=\"glossary-link\" title=\"Designing and testing model inputs to elicit useful behavior from a particular model and task.\" data-glossary-slug=\"prompt-engineering\">Prompt engineering<\/a> edits messages; context engineering owns the larger delivery pipeline and its lifecycle.<\/p>\n<h2>Check your understanding<\/h2>\n<p>More context is not automatically better context; relevance, timing, provenance, and hierarchy matter.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"How Claude Code Works in Large Codebases","url":"https:\/\/www.claude.com\/blog\/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start"},{"title":"BootstrapAgent: Distilling Repository Setup","url":"https:\/\/arxiv.org\/abs\/2605.15815"},{"title":"Why Software Factories Fail","url":"https:\/\/github.com\/humanlayer\/advanced-context-engineering-for-coding-agents\/blob\/main\/wsff.md"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/context-poisoning","slug":"context-poisoning","term":"Context poisoning","definition":"Corrupting information placed into an agent's active or persistent context so later decisions are based on false facts, malicious instructions, or distorted state.","definition_html":"<h2>Definition<\/h2>\n<p>Corrupting information placed into an agent's active or persistent context so later decisions are based on false facts, malicious instructions, or distorted state.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Context poisoning targets runtime or stored context; model poisoning alters model parameters or artifacts.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Persistent context needs validation, source attribution, correction, and deletion.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["memory poisoning"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"LLMs Corrupt Your Documents When You Delegate","url":"https:\/\/arxiv.org\/abs\/2604.15597"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/context-rot","slug":"context-rot","term":"Context rot","definition":"The degradation in an agent's ability to find, follow, or correctly weigh relevant information as its accumulated context becomes longer, noisier, stale, repetitive, or internally conflicting.","definition_html":"<h2>Definition<\/h2>\n<p>The degradation in an agent's ability to find, follow, or correctly weigh relevant information as its accumulated context becomes longer, noisier, stale, repetitive, or internally conflicting. Context rot can appear before the formal context-window limit because the problem is attention and instruction use, not merely storage capacity.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Context-window overflow means the input no longer fits. Context rot means the input fits but performs poorly. <a href=\"\/glossary\/context-poisoning\" class=\"glossary-link\" title=\"Corrupting information placed into an agent's active or persistent context so later decisions are based on false facts, malicious instructions, or distorted state.\" data-glossary-slug=\"context-poisoning\">Context poisoning<\/a> introduces malicious or misleading material; ordinary context rot can emerge from benign history, duplicated attempts, obsolete plans, and unresolved contradictions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A longer <a href=\"\/glossary\/context-window\" class=\"glossary-link\" title=\"The maximum token span a model can directly consider in one inference request, including instructions, conversation, retrieved material, tool schemas, and expected output.\" data-glossary-slug=\"context-window\">context window<\/a> can postpone truncation without preventing important instructions from being buried under irrelevant history.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["context degradation"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"context-window","url":"https:\/\/darkfactory.dev\/glossary\/context-window"},{"slug":"context-engineering","url":"https:\/\/darkfactory.dev\/glossary\/context-engineering"},{"slug":"prompt-compression","url":"https:\/\/darkfactory.dev\/glossary\/prompt-compression"},{"slug":"working-memory","url":"https:\/\/darkfactory.dev\/glossary\/working-memory"},{"slug":"context-poisoning","url":"https:\/\/darkfactory.dev\/glossary\/context-poisoning"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Long-Running Agents","url":"https:\/\/addyosmani.com\/blog\/long-running-agents\/"},{"title":"Instruction Adherence in Coding Agent Configuration Files","url":"https:\/\/arxiv.org\/abs\/2605.10039"},{"title":"The Anatomy of an Agent Harness","url":"https:\/\/www.langchain.com\/blog\/the-anatomy-of-an-agent-harness"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/context-window","slug":"context-window","term":"Context window","definition":"The maximum token span a model can directly consider in one inference request, including instructions, conversation, retrieved material, tool schemas, and expected output.","definition_html":"<h2>Definition<\/h2>\n<p>The maximum token span a model can directly consider in one inference request, including instructions, conversation, retrieved material, tool schemas, and expected output.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A large context window is capacity, not memory quality, instruction fidelity, or guaranteed recall.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Putting a fact in context does not prove it will influence the right decision.<\/p>\n","category":"foundations","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["context length"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"context-rot","url":"https:\/\/darkfactory.dev\/glossary\/context-rot"},{"slug":"prompt-compression","url":"https:\/\/darkfactory.dev\/glossary\/prompt-compression"},{"slug":"input-token","url":"https:\/\/darkfactory.dev\/glossary\/input-token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"tokenization-tax","url":"https:\/\/darkfactory.dev\/glossary\/tokenization-tax"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"How Claude Code Works in Large Codebases","url":"https:\/\/www.claude.com\/blog\/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/continuous-batching","slug":"continuous-batching","term":"Continuous batching","definition":"An inference scheduling technique that adds and removes generation requests at iteration boundaries as capacity becomes available.","definition_html":"<h2>Definition<\/h2>\n<p>An inference scheduling technique that adds and removes generation requests at iteration boundaries as capacity becomes available.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A static batch holds a fixed group of requests together. Continuous batching can admit a new request when another finishes, even while longer requests continue generating.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Continuous batching can improve accelerator use and throughput, but the scheduler still has to manage latency, fairness, and memory pressure.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["iteration-level scheduling"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"batch","url":"https:\/\/darkfactory.dev\/glossary\/batch"},{"slug":"latency","url":"https:\/\/darkfactory.dev\/glossary\/latency"},{"slug":"throughput","url":"https:\/\/darkfactory.dev\/glossary\/throughput"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Orca: A Distributed Serving System for Transformer-Based Generative Models","url":"https:\/\/www.usenix.org\/conference\/osdi22\/presentation\/yu"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/contrastive-learning","slug":"contrastive-learning","term":"Contrastive learning","definition":"A representation-learning approach that trains examples judged related to be closer and examples judged different to be farther apart in a learned space.","definition_html":"<h2>Definition<\/h2>\n<p>Contrastive learning is a representation-learning approach that trains examples judged related to be closer and examples judged different to be farther apart in a learned embedding space. Positive and negative relationships may come from labels, paired modalities, augmentations of the same example, or sampling rules.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Contrastive learning defines a training objective or strategy; an embedding is the resulting representation. It can be supervised or self-supervised and can align different modalities, such as images and text.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The definition of positive and negative pairs encodes assumptions. Poor pairing can teach shortcuts, false separation, or biased similarity even when the loss improves.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"embedding","url":"https:\/\/darkfactory.dev\/glossary\/embedding"},{"slug":"self-supervised-learning","url":"https:\/\/darkfactory.dev\/glossary\/self-supervised-learning"},{"slug":"multimodal-model","url":"https:\/\/darkfactory.dev\/glossary\/multimodal-model"},{"slug":"loss-function","url":"https:\/\/darkfactory.dev\/glossary\/loss-function"}],"related_factory_areas":[],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/control-graph","slug":"control-graph","term":"Control graph","definition":"A directed representation of the steps an agent system may execute and the conditions that select what runs next.","definition_html":"<h2>Definition<\/h2>\n<p>A directed representation of the steps an agent system may execute and the conditions that select what runs next. Nodes perform work; edges encode fixed, conditional, or dynamically created transitions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A control graph defines permitted execution structure. A <a href=\"\/glossary\/knowledge-graph\" class=\"glossary-link\" title=\"A structured representation of entities, concepts, and claims connected by named relationships and provenance.\" data-glossary-slug=\"knowledge-graph\">knowledge graph<\/a> represents entities and relationships. An <a href=\"\/glossary\/trace\" class=\"glossary-link\" title=\"A captured sequence of model calls, tool calls, events, timings, state changes, and outputs from an execution.\" data-glossary-slug=\"trace\">execution trace<\/a> records the path a particular run actually took.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Mark which transitions are enforced by code, which are chosen by a model, and which require an external signal or human decision.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["agent control graph"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"},{"slug":"orchestration","url":"https:\/\/darkfactory.dev\/glossary\/orchestration"},{"slug":"state-machine","url":"https:\/\/darkfactory.dev\/glossary\/state-machine"},{"slug":"directed-acyclic-graph","url":"https:\/\/darkfactory.dev\/glossary\/directed-acyclic-graph"},{"slug":"agent-loop","url":"https:\/\/darkfactory.dev\/glossary\/agent-loop"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"},{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/controlled-self-improvement","slug":"controlled-self-improvement","term":"Controlled self-improvement","definition":"Versioned modification of prompts, skills, memory, workflows, routing, or harness code under fixed evaluations, limited rollout, observation, and automatic reversion.","definition_html":"<h2>Definition<\/h2>\n<p>Versioned modification of prompts, skills, memory, workflows, routing, or harness code under fixed evaluations, limited rollout, observation, and automatic reversion.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>This changes the system around the model; recursive self-improvement is the broader and more speculative idea of an AI improving its own general capabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Keep the evaluator and authority boundary outside the component being mutated.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["harness self-improvement","self-improvement loop","learning loop","hill-climbing loop","hill climbing loop"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"improvement-graph","url":"https:\/\/darkfactory.dev\/glossary\/improvement-graph"},{"slug":"loop-engineering","url":"https:\/\/darkfactory.dev\/glossary\/loop-engineering"},{"slug":"verification-gate","url":"https:\/\/darkfactory.dev\/glossary\/verification-gate"},{"slug":"independent-verification","url":"https:\/\/darkfactory.dev\/glossary\/independent-verification"}],"related_factory_areas":[{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"Harness Engineering for Self-Improvement","url":"https:\/\/lilianweng.github.io\/posts\/2026-07-04-harness\/"},{"title":"Autoresearch as a Production Loop","url":"https:\/\/shopify.engineering\/autoresearch"},{"title":"The Therapist Pattern","url":"https:\/\/blog.fsck.com\/2026\/07\/20\/the-therapist-pattern\/"},{"title":"Building self-improving tax agents with Codex","url":"https:\/\/openai.com\/index\/building-self-improving-tax-agents-with-codex\/"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/convolutional-neural-network","slug":"convolutional-neural-network","term":"Convolutional neural network (CNN)","definition":"A neural-network architecture that applies learned local filters across spatial or sequential data.","definition_html":"<h2>Definition<\/h2>\n<p>A convolutional neural network applies shared learned filters across local regions, making it effective for spatial or sequential patterns.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A transformer uses attention to relate positions broadly; a CNN emphasizes local patterns and translation-related structure.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify the receptive field, shared filter, and output feature map in a simple image layer.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome","slug":"cost-per-accepted-durable-outcome","term":"Cost per accepted durable outcome","definition":"The total model, infrastructure, validation, retry, review, incident, and human-attention cost divided by outcomes that are accepted and remain useful over time.","definition_html":"<h2>Definition<\/h2>\n<p>The total model, infrastructure, validation, retry, review, incident, and human-attention cost divided by outcomes that are accepted and remain useful over time.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Token cost and task-success rate are inputs; this metric charges for failed attempts, downstream correction, and durability.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Define acceptance and durability windows before comparing factories or models.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["cost per accepted task","cost per successful task","cost per accepted result"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"},{"slug":"useful-intelligence-per-dollar","url":"https:\/\/darkfactory.dev\/glossary\/useful-intelligence-per-dollar"}],"related_factory_areas":[{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"GenAI Productivity and Learning: A Meta-Analysis","url":"https:\/\/arxiv.org\/abs\/2605.04779"},{"title":"The Economic Benefit of Refactoring","url":"https:\/\/martinfowler.com\/articles\/exploring-gen-ai\/refactoring-economic-benefit.html"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"},{"title":"A scorecard for the AI age","url":"https:\/\/openai.com\/index\/a-scorecard-for-the-ai-age\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/cross-entropy","slug":"cross-entropy","term":"Cross-entropy","definition":"A loss that measures how poorly a predicted probability distribution represents the target distribution.","definition_html":"<h2>Definition<\/h2>\n<p>A loss that measures how poorly a predicted <a href=\"\/glossary\/probability-distribution\" class=\"glossary-link\" title=\"A set of possible outcomes paired with nonnegative probabilities that sum to one.\" data-glossary-slug=\"probability-distribution\">probability distribution<\/a> represents the target distribution.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Cross-entropy is a specific loss over probability distributions. Accuracy counts correct decisions after choosing outputs and does not measure the quality of the full predicted distribution.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Confident probability assigned to the wrong class produces a larger penalty than a less confident mistake.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["cross-entropy loss"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"loss-function","url":"https:\/\/darkfactory.dev\/glossary\/loss-function"},{"slug":"probability-distribution","url":"https:\/\/darkfactory.dev\/glossary\/probability-distribution"},{"slug":"training","url":"https:\/\/darkfactory.dev\/glossary\/training"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/cross-validation","slug":"cross-validation","term":"Cross-validation","definition":"A resampling method that estimates generalization by repeatedly training and evaluating on different non-overlapping subsets of available data.","definition_html":"<h2>Definition<\/h2>\n<p>A resampling method that estimates generalization by repeatedly training and evaluating on different non-overlapping subsets of available data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A single train-test split produces one estimate. Cross-validation rotates which subset is held out, but it does not replace a final untouched <a href=\"\/glossary\/held-out-set\" class=\"glossary-link\" title=\"Evaluation examples deliberately withheld from training, prompt tuning, workflow design, or agent feedback to reduce leakage and overfitting.\" data-glossary-slug=\"held-out-set\">test set<\/a> when one is needed.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Splits must respect time, identity, and dependency boundaries or leakage can make the estimate falsely optimistic.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"generalization","url":"https:\/\/darkfactory.dev\/glossary\/generalization"},{"slug":"held-out-set","url":"https:\/\/darkfactory.dev\/glossary\/held-out-set"},{"slug":"overfitting","url":"https:\/\/darkfactory.dev\/glossary\/overfitting"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/dark-software-factory","slug":"dark-software-factory","term":"Dark software factory","definition":"A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.","definition_html":"<h2>Definition<\/h2>\n<p>A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A dark factory is not simply code written by AI or a <a href=\"\/glossary\/coding-agent\" class=\"glossary-link\" title=\"An AI agent equipped to inspect a software project, edit files, run development tools, test changes, and return or promote a software outcome.\" data-glossary-slug=\"coding-agent\">coding agent<\/a> with autonomy enabled; it is an assurance claim about the complete production loop.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask what domain is dark, who owns consequences, what evidence permits promotion, and how rollback and escalation work.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":true,"search_index_reason":"Owns the central reader question for this publication and distinguishes a full assurance-backed production loop from AI-assisted coding.","search_reviewed_at":"2026-08-06","aliases":["dark factory","lights-out factory","lights-out software factory","lights-off factory","lights-off software factory"],"link_forms":["dark software factories","dark factories","lights-out factories","lights-out software factories","lights-off factories","lights-off software factories"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"software-factory","url":"https:\/\/darkfactory.dev\/glossary\/software-factory"},{"slug":"agentic-software-engineering","url":"https:\/\/darkfactory.dev\/glossary\/agentic-software-engineering"},{"slug":"human-out-of-the-loop","url":"https:\/\/darkfactory.dev\/glossary\/human-out-of-the-loop"},{"slug":"risk-scoped-autonomy","url":"https:\/\/darkfactory.dev\/glossary\/risk-scoped-autonomy"}],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"StrongDM: Software Factories and the Agentic Moment","url":"https:\/\/factory.strongdm.ai\/"},{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"},{"title":"Why Software Factories Fail","url":"https:\/\/github.com\/humanlayer\/advanced-context-engineering-for-coding-agents\/blob\/main\/wsff.md"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/data-augmentation","slug":"data-augmentation","term":"Data augmentation","definition":"Expanding or varying training examples through transformations or generation intended to preserve task-relevant meaning.","definition_html":"<h2>Definition<\/h2>\n<p>Data augmentation expands or varies a training set by transforming existing examples or generating new ones while intending to preserve task-relevant meaning. Examples include cropping or rotating images, perturbing audio, paraphrasing text, generating counterexamples, and simulating rare conditions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/synthetic-data\" class=\"glossary-link\" title=\"Artificially generated records intended to reproduce useful properties of real data for training, testing, simulation, or privacy.\" data-glossary-slug=\"synthetic-data\">Synthetic data<\/a> can be created independently from simulations or generative systems; augmentation begins with a training objective and adds controlled variation. A transformation is harmful when it changes the correct label, erases important minority cases, or amplifies artifacts.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Validate augmented data against the invariance being assumed. More examples do not help when the transformation teaches the wrong task.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"training-data","url":"https:\/\/darkfactory.dev\/glossary\/training-data"},{"slug":"synthetic-data","url":"https:\/\/darkfactory.dev\/glossary\/synthetic-data"},{"slug":"generalization","url":"https:\/\/darkfactory.dev\/glossary\/generalization"},{"slug":"overfitting","url":"https:\/\/darkfactory.dev\/glossary\/overfitting"}],"related_factory_areas":[],"evidence":[{"title":"Andreessen Horowitz AI Glossary","url":"https:\/\/a16z.com\/ai-glossary\/"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/data-drift","slug":"data-drift","term":"Data drift","definition":"A change over time in the distribution of system inputs or features.","definition_html":"<h2>Definition<\/h2>\n<p>Data drift is a change over time in the distribution of inputs or features seen by a system.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/concept-drift\" class=\"glossary-link\" title=\"A change over time in the relationship between inputs and the correct target or decision.\" data-glossary-slug=\"concept-drift\">Concept drift<\/a> changes the relationship between inputs and desired outputs; data drift can occur without that relationship changing.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compare current feature distributions with the reference period and determine whether performance also changed.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/data-poisoning","slug":"data-poisoning","term":"Data poisoning","definition":"Introducing malicious, misleading, or strategically biased data into training, tuning, retrieval, memory, or evaluation pipelines to alter later behavior.","definition_html":"<h2>Definition<\/h2>\n<p>Introducing malicious, misleading, or strategically biased data into training, tuning, retrieval, memory, or evaluation pipelines to alter later behavior.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Training-data poisoning changes learned behavior; retrieval or <a href=\"\/glossary\/context-poisoning\" class=\"glossary-link\" title=\"Corrupting information placed into an agent's active or persistent context so later decisions are based on false facts, malicious instructions, or distorted state.\" data-glossary-slug=\"context-poisoning\">memory poisoning<\/a> changes runtime context without retraining.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Track provenance and mutation rights across every data plane, not only the training corpus.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["training-data poisoning"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"data-lifecycle","url":"https:\/\/darkfactory.dev\/factory\/data-lifecycle"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"LLMs Corrupt Your Documents When You Delegate","url":"https:\/\/arxiv.org\/abs\/2604.15597"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/dataset","slug":"dataset","term":"Dataset","definition":"A deliberately assembled collection of examples or records used to train, tune, evaluate, or operate an AI system.","definition_html":"<h2>Definition<\/h2>\n<p>A deliberately assembled collection of examples or records used to train, tune, evaluate, or operate an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A dataset is not automatically a training set; its role depends on how it is partitioned and used.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask who collected it, for what purpose, under what rights, and whether evaluation examples leaked into training.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["data set"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"data-lifecycle","url":"https:\/\/darkfactory.dev\/factory\/data-lifecycle"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/decoder","slug":"decoder","term":"Decoder","definition":"A component that transforms an internal representation or prior outputs into a target output sequence or reconstruction.","definition_html":"<h2>Definition<\/h2>\n<p>A decoder transforms an internal representation or prior outputs into a target output, reconstruction, or next-token sequence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An encoder maps input into representations; a decoder maps representations or preceding tokens toward outputs.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain what conditions the decoder receives and how it chooses the next output.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/decoder-only-model","slug":"decoder-only-model","term":"Decoder-only model","definition":"A sequence model that generates tokens causally from preceding context without a separate encoder component.","definition_html":"<h2>Definition<\/h2>\n<p>A decoder-only model generates each token conditioned on preceding tokens in one causal sequence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An <a href=\"\/glossary\/encoder-decoder-model\" class=\"glossary-link\" title=\"A sequence model in which an encoder represents the input and a decoder generates output conditioned on that representation.\" data-glossary-slug=\"encoder-decoder-model\">encoder-decoder model<\/a> separately encodes source input; a decoder-only model places instructions, context, and output into one continuation format.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Show how prompt and response become one sequence while masking future tokens during training.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/deep-learning","slug":"deep-learning","term":"Deep learning","definition":"Machine learning based on neural networks with multiple representational layers, allowing complex features to be learned from data.","definition_html":"<h2>Definition<\/h2>\n<p><a href=\"\/glossary\/machine-learning\" class=\"glossary-link\" title=\"A family of methods in which computational models improve task performance by finding patterns in data rather than relying only on explicitly programmed rules.\" data-glossary-slug=\"machine-learning\">Machine learning<\/a> based on neural networks with multiple representational layers, allowing complex features to be learned from data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Deep learning is a subset of machine learning; not every ML model is a deep <a href=\"\/glossary\/neural-network\" class=\"glossary-link\" title=\"A parameterized computational model composed of connected layers that transform representations and learn by adjusting weights to reduce an objective.\" data-glossary-slug=\"neural-network\">neural network<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Linear regression and decision trees are machine learning but not deep learning.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/differential-privacy","slug":"differential-privacy","term":"Differential privacy","definition":"A mathematical privacy framework that bounds how much a computation's output can change because one person's data is included or removed.","definition_html":"<h2>Definition<\/h2>\n<p>Differential privacy provides a mathematical bound on how much a computation's output can change when one person's record is added or removed.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>De-identification removes or masks identifiers; differential privacy limits inferential influence under a quantified privacy budget.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the privacy parameters, unit of protection, composition across queries, and utility cost.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/diffusion-model","slug":"diffusion-model","term":"Diffusion model","definition":"A generative model that learns to reverse a gradual noising process to create data from noise.","definition_html":"<h2>Definition<\/h2>\n<p>A diffusion model learns a denoising process that transforms noise into samples resembling its training distribution.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A GAN generates through an adversarially trained generator; diffusion typically generates through repeated denoising steps.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Describe the forward noising process and the learned reverse process.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/digital-twin","slug":"digital-twin","term":"Digital twin","definition":"A sufficiently faithful executable representation of a system or environment used to test behavior, scenarios, or changes before affecting the real target.","definition_html":"<h2>Definition<\/h2>\n<p>A sufficiently faithful executable representation of a system or environment used to test behavior, scenarios, or changes before affecting the real target.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A digital twin models operational behavior; a test fixture may cover only a narrow example or component.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Twin value depends on fidelity, coverage, freshness, and explicit knowledge of where the model diverges from reality.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"StrongDM: Software Factories and the Agentic Moment","url":"https:\/\/factory.strongdm.ai\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/direct-preference-optimization","slug":"direct-preference-optimization","term":"Direct preference optimization (DPO)","definition":"A preference-training method that directly adjusts a model toward preferred responses and away from rejected ones without first training a separate reward model in the classic RLHF pipeline.","definition_html":"<h2>Definition<\/h2>\n<p>A preference-training method that directly adjusts a model toward preferred responses and away from rejected ones without first training a separate reward model in the classic RLHF pipeline.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>DPO is one model-training technique, not a general term for evaluating preferences.<\/p>\n<h2>Check your understanding<\/h2>\n<p>It changes <a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a> and should not be confused with runtime ranking of candidate outputs.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["DPO"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/direct-prompt-injection","slug":"direct-prompt-injection","term":"Direct prompt injection","definition":"Prompt injection delivered directly through the current user's message or another explicit input channel.","definition_html":"<h2>Definition<\/h2>\n<p><a href=\"\/glossary\/prompt-injection\" class=\"glossary-link\" title=\"Manipulating an AI system by placing instructions in input or data that the model treats as authoritative enough to alter intended behavior.\" data-glossary-slug=\"prompt-injection\">Prompt injection<\/a> delivered directly through the current user's message or another explicit input channel.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Direct injection comes from the interacting user; indirect injection arrives through content the system retrieves or observes.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Authorized users can still be attackers, so authentication alone does not prevent direct injection.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/directed-acyclic-graph","slug":"directed-acyclic-graph","term":"Directed acyclic graph (DAG)","definition":"A directed graph with no path that returns to an earlier node.","definition_html":"<h2>Definition<\/h2>\n<p>A directed graph with no path that returns to an earlier node. DAGs naturally represent one-way dependency structures and work that must eventually complete without revisiting a prior state.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Many production agent workflows are not DAGs because they retry, revise, ask for missing information, call tools repeatedly, or pause and resume. Those behaviors introduce cycles or external state transitions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If a workflow may revisit a prior node, it is not acyclic even if its happy path is drawn left to right.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["DAG"],"link_forms":["directed acyclic graphs","DAGs"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"},{"slug":"agent-loop","url":"https:\/\/darkfactory.dev\/glossary\/agent-loop"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/durable-memory","slug":"durable-memory","term":"Durable memory","definition":"State intentionally retained across runs, such as verified facts, decisions, preferences, learned procedures, or persistent identity.","definition_html":"<h2>Definition<\/h2>\n<p>State intentionally retained across runs, such as verified facts, decisions, preferences, learned procedures, or persistent identity.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Durable memory is an application artifact, not knowledge permanently learned into <a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Persistent memory needs ownership, provenance, correction, deletion, and mutation controls.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["long-term memory"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"BootstrapAgent: Distilling Repository Setup","url":"https:\/\/arxiv.org\/abs\/2605.15815"},{"title":"The Therapist Pattern","url":"https:\/\/blog.fsck.com\/2026\/07\/20\/the-therapist-pattern\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/embedding","slug":"embedding","term":"Embedding","definition":"A learned numeric vector that represents an item so that useful semantic or structural relationships can be measured geometrically.","definition_html":"<h2>Definition<\/h2>\n<p>A learned numeric vector that represents an item so that useful semantic or structural relationships can be measured geometrically.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An embedding is the representation; a <a href=\"\/glossary\/vector-database\" class=\"glossary-link\" title=\"A data system designed to store embeddings and retrieve items by vector similarity, often with metadata filtering.\" data-glossary-slug=\"vector-database\">vector database<\/a> stores and searches vectors; <a href=\"\/glossary\/semantic-search\" class=\"glossary-link\" title=\"Retrieval based primarily on meaning represented by embeddings rather than exact keyword overlap.\" data-glossary-slug=\"semantic-search\">semantic search<\/a> is an application of them.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Nearby vectors indicate similarity under the embedding model, not factual equivalence or truth.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["vector embedding"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/emergent-behavior","slug":"emergent-behavior","term":"Emergent behavior","definition":"A capability or behavior that appears qualitatively new at a larger scale or higher level of system interaction rather than as an obvious continuation of smaller-scale measurements.","definition_html":"<h2>Definition<\/h2>\n<p>A capability or behavior that appears qualitatively new at a larger scale or higher level of system interaction rather than as an obvious continuation of smaller-scale measurements.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Some apparent emergence reflects a real system-level threshold. Some reflects a discontinuous benchmark metric applied to smoothly changing performance. The observation and the explanation must be evaluated separately.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Before calling a behavior emergent, inspect the raw measurements, metric shape, model scales tested, and whether interaction effects could explain the threshold.<\/p>\n","category":"evaluation-and-reliability","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["emergent ability"],"link_forms":["emergent behaviors","emergent abilities"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"benchmark","url":"https:\/\/darkfactory.dev\/glossary\/benchmark"},{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"generalization","url":"https:\/\/darkfactory.dev\/glossary\/generalization"}],"related_factory_areas":[],"evidence":[{"title":"Emergent Abilities of Large Language Models","url":"https:\/\/arxiv.org\/abs\/2206.07682"},{"title":"Are Emergent Abilities of Large Language Models a Mirage?","url":"https:\/\/arxiv.org\/abs\/2304.15004"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/encoder","slug":"encoder","term":"Encoder","definition":"A component that transforms input into an internal representation useful for later prediction or generation.","definition_html":"<h2>Definition<\/h2>\n<p>An encoder transforms input into an internal representation useful for later prediction, retrieval, classification, or generation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A decoder turns a representation or prior tokens into outputs; an encoder creates the representation it consumes.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify the input, representation, and downstream task the encoding must preserve.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/encoder-decoder-model","slug":"encoder-decoder-model","term":"Encoder-decoder model","definition":"A sequence model in which an encoder represents the input and a decoder generates output conditioned on that representation.","definition_html":"<h2>Definition<\/h2>\n<p>An encoder-decoder model first represents an input with an encoder, then generates a target output with a decoder conditioned on that representation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/decoder-only-model\" class=\"glossary-link\" title=\"A sequence model that generates tokens causally from preceding context without a separate encoder component.\" data-glossary-slug=\"decoder-only-model\">decoder-only model<\/a> predicts continuations from a single causal sequence rather than using a separate input encoder.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Map the source input, encoder representation, decoder context, and target output in translation.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/ephemeral-environment","slug":"ephemeral-environment","term":"Ephemeral environment","definition":"A short-lived execution environment created for a run and discarded afterward.","definition_html":"<h2>Definition<\/h2>\n<p>A short-lived execution environment created for a run and discarded afterward.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Ephemerality limits persistent residue but is not equivalent to isolation, clean identity, or safe reconstruction.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Fresh environments can also help an attacker repeatedly rebuild a foothold if external state and egress remain available.<\/p>\n","category":"tools-and-protocols","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["ephemeral sandbox"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"}],"evidence":[{"title":"Long-Running Agents","url":"https:\/\/addyosmani.com\/blog\/long-running-agents\/"},{"title":"Anatomy of a Frontier Lab Agent Intrusion","url":"https:\/\/huggingface.co\/blog\/agent-intrusion-technical-timeline"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/epoch","slug":"epoch","term":"Epoch","definition":"One complete pass through the training dataset, usually divided into batches.","definition_html":"<h2>Definition<\/h2>\n<p>An epoch is one complete pass through the <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training dataset<\/a>, usually divided into batches.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A training step usually processes one batch; an epoch consists of enough steps to cover the dataset once.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Given dataset and batch sizes, estimate the steps per epoch and explain why more epochs can overfit.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/evaluation","slug":"evaluation","term":"Evaluation (eval)","definition":"A systematic measurement of model or system behavior against defined tasks, criteria, datasets, or operational outcomes.","definition_html":"<h2>Definition<\/h2>\n<p>A systematic measurement of model or system behavior against defined tasks, criteria, datasets, or operational outcomes.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An evaluation is the measurement process; a benchmark is a standardized evaluation set or protocol.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The reported unit must say whether it measures a model, harness, full agent system, or production workflow.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["eval"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"AgentAtlas: Control-Decision Taxonomy","url":"https:\/\/arxiv.org\/abs\/2605.20530"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/excessive-agency","slug":"excessive-agency","term":"Excessive agency","definition":"Granting an AI system more functionality, permissions, autonomy, or action scope than required for its intended task.","definition_html":"<h2>Definition<\/h2>\n<p>Granting an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> more functionality, permissions, autonomy, or action scope than required for its intended task.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Agency is capacity to act; excessive agency is the avoidable security condition created by overbroad authority.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Reduce tools, privileges, scope, duration, and irreversible actions rather than relying only on better prompting.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/execution-graph","slug":"execution-graph","term":"Execution graph","definition":"A run-oriented graph of executable steps and the control or data dependencies connecting them.","definition_html":"<h2>Definition<\/h2>\n<p>A run-oriented graph of executable steps and the control or data dependencies connecting them. It may be partially known before execution and expanded as an orchestrator creates work at runtime.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/control-graph\" class=\"glossary-link\" title=\"A directed representation of the steps an agent system may execute and the conditions that select what runs next.\" data-glossary-slug=\"control-graph\">control graph<\/a> describes allowed routing; an execution graph represents the work instantiated for a run; an <a href=\"\/glossary\/trace\" class=\"glossary-link\" title=\"A captured sequence of model calls, tool calls, events, timings, state changes, and outputs from an execution.\" data-glossary-slug=\"trace\">execution trace<\/a> records the ordered events and outcomes that occurred.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The graph should reveal ownership, dependencies, concurrency limits, retry boundaries, and the join points at which results become usable.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["task execution graph"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"},{"slug":"orchestration","url":"https:\/\/darkfactory.dev\/glossary\/orchestration"},{"slug":"trace","url":"https:\/\/darkfactory.dev\/glossary\/trace"},{"slug":"execution-lineage","url":"https:\/\/darkfactory.dev\/glossary\/execution-lineage"},{"slug":"task-decomposition","url":"https:\/\/darkfactory.dev\/glossary\/task-decomposition"},{"slug":"directed-acyclic-graph","url":"https:\/\/darkfactory.dev\/glossary\/directed-acyclic-graph"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"}],"evidence":[{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"},{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"},{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/execution-lineage","slug":"execution-lineage","term":"Execution lineage","definition":"A reconstructable record linking an outcome to the initiating intent, model, harness, tools, inputs, actions, evidence, and state transitions that produced it.","definition_html":"<h2>Definition<\/h2>\n<p>A reconstructable record linking an outcome to the initiating intent, model, harness, tools, inputs, actions, evidence, and state transitions that produced it.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Lineage records causal provenance across execution; a log may contain events without enough structure to reconstruct causality.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A factory should be able to answer what acted, under which policy, using which evidence, and why promotion was permitted.<\/p>\n","category":"tools-and-protocols","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["lineage"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"trace","url":"https:\/\/darkfactory.dev\/glossary\/trace"}],"related_factory_areas":[{"slug":"governance-accountability","url":"https:\/\/darkfactory.dev\/factory\/governance-accountability"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Execution Lineage for Reproducible AI-Native Work","url":"https:\/\/arxiv.org\/abs\/2605.06365"},{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/expert-system","slug":"expert-system","term":"Expert system","definition":"A system that applies an explicitly represented knowledge base and inference rules to make recommendations or decisions within a bounded domain.","definition_html":"<h2>Definition<\/h2>\n<p>An expert system applies an explicitly represented knowledge base and inference rules to make recommendations or decisions within a bounded domain. Classic systems separate domain facts from an inference engine that evaluates rules such as if-then conditions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Expert systems encode much of their decision logic explicitly; machine-learning models learn statistical parameters from data. Modern systems can combine both, using rules or policy engines to constrain probabilistic models.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explicit rules can improve traceability but still fail when the knowledge base is incomplete, stale, contradictory, or applied outside its intended domain.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["expert systems"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"ai-system","url":"https:\/\/darkfactory.dev\/glossary\/ai-system"},{"slug":"narrow-ai","url":"https:\/\/darkfactory.dev\/glossary\/narrow-ai"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"explainability","url":"https:\/\/darkfactory.dev\/glossary\/explainability"}],"related_factory_areas":[],"evidence":[{"title":"Andreessen Horowitz AI Glossary","url":"https:\/\/a16z.com\/ai-glossary\/"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/explainability","slug":"explainability","term":"Explainability","definition":"The ability to provide a human-usable account of why a system produced a particular output or action.","definition_html":"<h2>Definition<\/h2>\n<p>Explainability is the ability to provide a human-usable account of why a system produced a particular output or action.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A rationale generated after the fact can sound plausible without faithfully reflecting model causation; explanation quality therefore needs independent validation.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask whether the explanation is faithful, useful to its audience, stable, and sufficient for the decision at stake.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/f1-score","slug":"f1-score","term":"F1 score","definition":"The harmonic mean of precision and recall.","definition_html":"<h2>Definition<\/h2>\n<p>The F1 score is the harmonic mean of precision and recall, producing one measure that is high only when both are high.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Accuracy includes true negatives; F1 ignores them and does not express asymmetric costs or calibration.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compute F1 from precision and recall, then state what important information the single number discards.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/false-negative","slug":"false-negative","term":"False negative","definition":"An outcome incorrectly classified as absent when it is actually present.","definition_html":"<h2>Definition<\/h2>\n<p>A false negative occurs when a system fails to identify a condition that is actually present.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/false-positive\" class=\"glossary-link\" title=\"An outcome incorrectly classified or flagged as present when it is actually absent.\" data-glossary-slug=\"false-positive\">false positive<\/a> raises an incorrect alarm; a false negative creates a miss.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the real-world consequence of a miss and whether delayed detection can recover.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/false-positive","slug":"false-positive","term":"False positive","definition":"An outcome incorrectly classified or flagged as present when it is actually absent.","definition_html":"<h2>Definition<\/h2>\n<p>A false positive occurs when a system predicts or flags a condition that is not actually present.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/false-negative\" class=\"glossary-link\" title=\"An outcome incorrectly classified as absent when it is actually present.\" data-glossary-slug=\"false-negative\">false negative<\/a> misses a condition that is present; their costs are domain-specific and rarely symmetric.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the real-world consequence of a false alarm and who bears it.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/feature","slug":"feature","term":"Feature","definition":"A measurable input attribute or derived representation used by a machine-learning system.","definition_html":"<h2>Definition<\/h2>\n<p>A feature is a measurable input attribute or derived representation supplied to a machine-learning model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A raw field is collected data; a feature is the representation actually used by the model, which may be transformed or learned.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask how the feature is computed, when it is available, and whether it leaks the target.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/federated-learning","slug":"federated-learning","term":"Federated learning","definition":"Training a shared model across distributed data holders without centralizing their raw training data.","definition_html":"<h2>Definition<\/h2>\n<p>Federated learning trains a shared model across distributed participants while keeping raw <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a> at their locations.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>It reduces raw-data centralization but does not automatically prevent information leakage from updates or solve participant trust.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Trace what leaves each participant, who aggregates it, and which privacy or security controls protect updates.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/few-shot-prompting","slug":"few-shot-prompting","term":"Few-shot prompting","definition":"Supplying a small set of worked examples in context to steer task behavior without updating model weights.","definition_html":"<h2>Definition<\/h2>\n<p>Few-shot prompting supplies several examples in the current context to demonstrate a task, format, or decision boundary.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Fine-tuning updates parameters across training examples; few-shot prompting conditions one inference request without weight updates.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Choose examples that cover boundaries and explain how order or imbalance might bias the result.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/fine-tuning","slug":"fine-tuning","term":"Fine-tuning","definition":"An additional training phase that adapts a pretrained model by updating some or all parameters using task- or domain-specific data.","definition_html":"<h2>Definition<\/h2>\n<p>An additional training phase that adapts a pretrained model by updating some or all parameters using task- or domain-specific data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Fine-tuning changes weights; prompting, retrieval, and tools change runtime inputs or capabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Fine-tuning may improve a target behavior while degrading others, so it still requires evaluation.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["finetuning","adaptation"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"low-rank-adaptation","url":"https:\/\/darkfactory.dev\/glossary\/low-rank-adaptation"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/foundation-model","slug":"foundation-model","term":"Foundation model","definition":"A broadly trained model, usually learned through self-supervision on diverse data, that can be adapted to many downstream tasks.","definition_html":"<h2>Definition<\/h2>\n<p>A broadly trained model, usually learned through self-supervision on diverse data, that can be adapted to many downstream tasks.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A foundation model is defined by broad pretraining and adaptability; an <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a> is the language-oriented subset most people encounter.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A fine-tuned application model may descend from a foundation model without itself serving as a broad foundation.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/frontier-model","slug":"frontier-model","term":"Frontier model","definition":"A general-purpose AI model at or near the leading edge of broadly evaluated capability at a particular time.","definition_html":"<h2>Definition<\/h2>\n<p>A frontier model is a general-purpose <a href=\"\/glossary\/ai-model\" class=\"glossary-link\" title=\"A computational component whose learned or encoded structure transforms inputs into outputs such as scores, predictions, classifications, or generated content.\" data-glossary-slug=\"ai-model\">AI model<\/a> at or near the leading edge of broadly evaluated capability at a particular time. The term identifies a model's position relative to a changing comparison set; it does not name a fixed architecture, parameter count, vendor class, license, or permanent tier.<\/p>\n<p>Government and international-safety sources commonly anchor the category to models that match or exceed the capabilities of the most advanced contemporary systems across a wide range of tasks. In everyday technical use, the boundary is looser: a model may be called frontier because it leads important evaluations, introduces consequential capabilities, or materially advances the practical state of the art.<\/p>\n<h2>What makes the category difficult<\/h2>\n<p>Frontier status is:<\/p>\n<ul>\n<li><strong>Relative:<\/strong> a model can leave the frontier as newer models improve.<\/li>\n<li><strong>Multidimensional:<\/strong> leadership in coding does not prove leadership in vision, reasoning, tool use, safety, latency, or cost.<\/li>\n<li><strong>Evaluation-dependent:<\/strong> rankings change with benchmarks, prompts, scaffolds, inference budgets, and contamination controls.<\/li>\n<li><strong>System-sensitive:<\/strong> the same base model can perform differently when paired with different tools, context, retrieval, or agent harnesses.<\/li>\n<li><strong>Purpose-sensitive:<\/strong> policy discussions often use the term to identify models requiring enhanced evaluation or safeguards, while product discussions may use it simply to mean premium or state of the art.<\/li>\n<\/ul>\n<p>There is no universal score or compute threshold that permanently determines frontier status. Any serious claim should therefore state the date, capability domain, evaluation method, and comparison set.<\/p>\n<h2>Why it matters in a software factory<\/h2>\n<p>Frontier models may expand the set of tasks that can be delegated, but capability alone does not make them the correct default. They can carry higher cost, latency, rate-limit exposure, vendor dependency, nondeterminism, and operational <a href=\"\/glossary\/blast-radius\" class=\"glossary-link\" title=\"The maximum plausible scope of harm, data exposure, or irreversible change if an action or component fails or is compromised.\" data-glossary-slug=\"blast-radius\">blast radius<\/a>. They may also fail differently from smaller or older models.<\/p>\n<p>Model routing should select the least costly model that satisfies the task's <a href=\"\/glossary\/acceptance-criteria\" class=\"glossary-link\" title=\"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.\" data-glossary-slug=\"acceptance-criteria\">acceptance criteria<\/a> and risk constraints. A frontier label is evidence that a model deserves evaluation, not evidence that its output deserves acceptance.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<ul>\n<li>A <strong><a href=\"\/glossary\/foundation-model\" class=\"glossary-link\" title=\"A broadly trained model, usually learned through self-supervision on diverse data, that can be adapted to many downstream tasks.\" data-glossary-slug=\"foundation-model\">foundation model<\/a><\/strong> is broadly pretrained and adaptable. Many foundation models are not frontier models.<\/li>\n<li>A <strong><a href=\"\/glossary\/reasoning-model\" class=\"glossary-link\" title=\"A model optimized to spend additional inference effort on multi-step problem solving before returning an answer or action.\" data-glossary-slug=\"reasoning-model\">reasoning model<\/a><\/strong> uses training or inference techniques optimized for multi-step problem solving. It may or may not be frontier overall.<\/li>\n<li>An <strong><a href=\"\/glossary\/open-weight-model\" class=\"glossary-link\" title=\"A model whose trained parameters are available to download, inspect, or run under stated terms, without implying that the complete system is open source.\" data-glossary-slug=\"open-weight-model\">open-weight model<\/a><\/strong> exposes trained parameters. Openness and capability position are independent axes.<\/li>\n<li>A <strong><a href=\"\/glossary\/proprietary-model\" class=\"glossary-link\" title=\"A model whose weights, development artifacts, or rights to inspect, modify, run, or redistribute it remain materially controlled by an owner.\" data-glossary-slug=\"proprietary-model\">proprietary model<\/a><\/strong> restricts artifacts or usage rights. Many frontier models are proprietary, but the words are not synonyms.<\/li>\n<li><strong><a href=\"\/glossary\/artificial-general-intelligence\" class=\"glossary-link\" title=\"A contested term for AI with broad, transferable competence across many cognitive tasks, often at or beyond human-level breadth.\" data-glossary-slug=\"artificial-general-intelligence\">Artificial general intelligence<\/a><\/strong> is a disputed capability threshold or aspiration. Frontier describes the leading edge that exists now, not a claim that AGI has been reached.<\/li>\n<\/ul>\n<h2>Check your understanding<\/h2>\n<p>\"Frontier\" is incomplete without a date and a capability claim. Ask: frontier at what, measured how, against which models, and under what harness and <a href=\"\/glossary\/token-budget\" class=\"glossary-link\" title=\"An explicit allocation or ceiling for token consumption across a request, run, task, user, workflow, or time period.\" data-glossary-slug=\"token-budget\">inference budget<\/a>?<\/p>\n","category":"models-and-training","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["frontier AI model"],"link_forms":["frontier models","frontier AI models"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"foundation-model","url":"https:\/\/darkfactory.dev\/glossary\/foundation-model"},{"slug":"benchmark","url":"https:\/\/darkfactory.dev\/glossary\/benchmark"},{"slug":"capability","url":"https:\/\/darkfactory.dev\/glossary\/capability"},{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"open-weight-model","url":"https:\/\/darkfactory.dev\/glossary\/open-weight-model"},{"slug":"proprietary-model","url":"https:\/\/darkfactory.dev\/glossary\/proprietary-model"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"UK AI Safety Summit: What Is Frontier AI?","url":"https:\/\/www.gov.uk\/government\/publications\/ai-safety-summit-introduction\/ai-safety-summit-introduction-html"},{"title":"International AI Safety Report 2026","url":"https:\/\/internationalaisafetyreport.org\/publication\/international-ai-safety-report-2026"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/function-calling","slug":"function-calling","term":"Function calling","definition":"A model interface in which the model selects a named function and supplies structured arguments for application code to execute.","definition_html":"<h2>Definition<\/h2>\n<p>A model interface in which the model selects a named function and supplies structured arguments for application code to execute.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The model proposes a call; the application remains responsible for authorization, validation, execution, and error handling.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A function call is not proof that the requested action is permitted or correctly bound to user intent.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["tool calling"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"},{"title":"Deterministic Tool-Schema Compilation","url":"https:\/\/arxiv.org\/abs\/2605.04107"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/generalization","slug":"generalization","term":"Generalization","definition":"The ability of a learned model or system to perform well on relevant examples, tasks, or environments not used to fit it.","definition_html":"<h2>Definition<\/h2>\n<p>The ability of a learned model or system to perform well on relevant examples, tasks, or environments not used to fit it.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Generalization concerns transfer beyond fitted examples; benchmark performance measures only one sampled regime.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A system may generalize within a benchmark family but fail under changed tools, budgets, or production conditions.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"SpecBench: the reward-hacking gap grows with codebase size","url":"https:\/\/arxiv.org\/abs\/2605.21384"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/generative-ai","slug":"generative-ai","term":"Generative AI","definition":"AI designed to produce new content, such as text, code, images, audio, video, or structured data, based on patterns learned from data.","definition_html":"<h2>Definition<\/h2>\n<p>AI designed to produce new content, such as text, code, images, audio, video, or structured data, based on patterns learned from data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Generative AI describes the kind of output, not whether a system is autonomous or agentic.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A chatbot can be generative without being an agent; an agent may use generative models while also acting through tools.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["GenAI","generative artificial intelligence"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/generative-adversarial-network","slug":"generative-adversarial-network","term":"Generative adversarial network (GAN)","definition":"A generative architecture trained through competition between a generator and a discriminator.","definition_html":"<h2>Definition<\/h2>\n<p>A generative adversarial network trains a generator to produce samples and a discriminator to distinguish generated samples from real ones.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/diffusion-model\" class=\"glossary-link\" title=\"A generative model that learns to reverse a gradual noising process to create data from noise.\" data-glossary-slug=\"diffusion-model\">diffusion model<\/a> learns to reverse a noise process; a GAN learns through an adversarial game.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State what each network optimizes and what happens if one becomes much stronger than the other.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/golden-set","slug":"golden-set","term":"Golden set","definition":"A curated collection of examples with trusted expected outcomes used for regression testing or evaluation.","definition_html":"<h2>Definition<\/h2>\n<p>A curated collection of examples with trusted expected outcomes used for regression testing or evaluation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A golden set is trusted reference data; a <a href=\"\/glossary\/held-out-set\" class=\"glossary-link\" title=\"Evaluation examples deliberately withheld from training, prompt tuning, workflow design, or agent feedback to reduce leakage and overfitting.\" data-glossary-slug=\"held-out-set\">held-out set<\/a> is defined by separation from tuning, whether or not labels are perfect.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Golden sets require maintenance when reality, policy, or product behavior changes.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["gold set","golden dataset"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/grader","slug":"grader","term":"Grader","definition":"A component that scores, classifies, or judges an output or trajectory against a rubric or expected behavior.","definition_html":"<h2>Definition<\/h2>\n<p>A component that scores, classifies, or judges an output or trajectory against a rubric or expected behavior.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A grader produces a judgment; an oracle establishes ground truth for a property.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Graders need calibration, abstention, adversarial testing, and protection from consequence-sensitive framing.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["evaluator","critic"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"verification-loop","url":"https:\/\/darkfactory.dev\/glossary\/verification-loop"},{"slug":"oracle","url":"https:\/\/darkfactory.dev\/glossary\/oracle"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"AgentAtlas: Control-Decision Taxonomy","url":"https:\/\/arxiv.org\/abs\/2605.20530"},{"title":"Agentic Misalignment in Summer 2026","url":"https:\/\/alignment.anthropic.com\/2026\/agentic-misalignment-summer-2026\/"},{"title":"Where Does Agent Reliability Come From?","url":"https:\/\/arxiv.org\/abs\/2607.17044"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/gradient-descent","slug":"gradient-descent","term":"Gradient descent","definition":"An optimization method that iteratively changes parameters in the direction expected to reduce loss.","definition_html":"<h2>Definition<\/h2>\n<p>Gradient descent iteratively changes parameters using loss gradients so the objective is expected to improve.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Backpropagation computes gradients through a <a href=\"\/glossary\/neural-network\" class=\"glossary-link\" title=\"A parameterized computational model composed of connected layers that transform representations and learn by adjusting weights to reduce an objective.\" data-glossary-slug=\"neural-network\">neural network<\/a>; gradient descent uses those gradients to update parameters.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Describe the current parameters, computed gradient, step size, and update.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/graph-engineering","slug":"graph-engineering","term":"Graph engineering","definition":"Designing an agent system as explicit nodes, state, and transitions so deterministic control and model judgment have visible boundaries.","definition_html":"<h2>Definition<\/h2>\n<p>Designing an agent system as explicit nodes, state, and transitions so deterministic control and model judgment have visible boundaries. A node may contain ordinary code, a model call, a tool, or a complete agent; an edge describes what may happen next.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Graph engineering is a contested umbrella term, not a settled new discipline. It can refer to control flow, knowledge representation, execution structure, or improvement relationships, which should be named separately. <a href=\"\/glossary\/loop-engineering\" class=\"glossary-link\" title=\"Designing the feedback cycles around an agent so action, verification, operational triggers, and system improvement have explicit state, evidence, limits, and stopping conditions.\" data-glossary-slug=\"loop-engineering\">Loop engineering<\/a> is not its opposite because a loop is a cyclic graph.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A graph earns its complexity only when it makes an important path, boundary, state transition, retry, or gate easier to enforce and inspect.<\/p>\n","category":"agents-and-automation","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"improvement-graph","url":"https:\/\/darkfactory.dev\/glossary\/improvement-graph"},{"slug":"agent-loop","url":"https:\/\/darkfactory.dev\/glossary\/agent-loop"},{"slug":"loop-engineering","url":"https:\/\/darkfactory.dev\/glossary\/loop-engineering"},{"slug":"orchestration","url":"https:\/\/darkfactory.dev\/glossary\/orchestration"},{"slug":"state-machine","url":"https:\/\/darkfactory.dev\/glossary\/state-machine"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"},{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"},{"title":"Bouchard: Graph Engineering Explained","url":"https:\/\/www.louisbouchard.ai\/graph-engineering-explained\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/graphrag","slug":"graphrag","term":"GraphRAG","definition":"Retrieval-augmented generation that builds and queries graph structure, often alongside source text and vector retrieval.","definition_html":"<h2>Definition<\/h2>\n<p><a href=\"\/glossary\/retrieval-augmented-generation\" class=\"glossary-link\" title=\"Generating a response after retrieving relevant material from an external knowledge source and adding it to model context.\" data-glossary-slug=\"retrieval-augmented-generation\">Retrieval-augmented generation<\/a> that builds and queries graph structure, often alongside source text and vector retrieval. Microsoft GraphRAG extracts entities, relationships, claims, communities, and summaries, then uses those derived structures during retrieval.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>GraphRAG is a form of RAG, not its replacement. It is useful when a question depends on relationships or corpus-level themes that similarity search may miss, but it adds indexing cost and model-extraction error.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Evaluate GraphRAG against a simpler retrieval baseline on the actual question class, and preserve links from every generated structure back to source text.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["graph retrieval-augmented generation","graph-based RAG"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"retrieval-augmented-generation","url":"https:\/\/darkfactory.dev\/glossary\/retrieval-augmented-generation"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"vector-database","url":"https:\/\/darkfactory.dev\/glossary\/vector-database"},{"slug":"semantic-search","url":"https:\/\/darkfactory.dev\/glossary\/semantic-search"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Microsoft GraphRAG Documentation","url":"https:\/\/microsoft.github.io\/graphrag\/"},{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/gpu","slug":"gpu","term":"Graphics processing unit (GPU)","definition":"A highly parallel processor widely used to train and run neural networks.","definition_html":"<h2>Definition<\/h2>\n<p>A graphics processing unit is a highly parallel processor widely used for the matrix operations involved in neural-network training and inference.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A model is software and learned parameters; a GPU is hardware that executes its computations.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain how memory capacity, memory bandwidth, precision, and batching affect usable model performance.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/greedy-decoding","slug":"greedy-decoding","term":"Greedy decoding","definition":"Generating each next token by selecting the current highest-probability candidate.","definition_html":"<h2>Definition<\/h2>\n<p>Greedy decoding selects the highest-probability token at every generation step.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Sampling draws from a distribution and <a href=\"\/glossary\/beam-search\" class=\"glossary-link\" title=\"A decoding algorithm that keeps a fixed number of high-scoring partial sequences while generating output.\" data-glossary-slug=\"beam-search\">beam search<\/a> retains multiple candidate sequences; greedy decoding keeps only one local choice.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain why the locally most likely token sequence need not be the globally most likely or best sequence.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/grounding","slug":"grounding","term":"Grounding","definition":"Connecting an AI output to identifiable evidence, data, observations, or constraints outside the model's unsupported generation.","definition_html":"<h2>Definition<\/h2>\n<p>Connecting an AI output to identifiable evidence, data, observations, or constraints outside the model's unsupported generation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Grounding supplies an evidentiary basis; retrieval merely supplies candidate context and may retrieve wrong or poisoned material.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask whether a claim can be traced to a source and whether that source actually entails it.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["grounded generation"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Theory Under Construction (Comet-H)","url":"https:\/\/arxiv.org\/abs\/2604.27209"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/guardrail","slug":"guardrail","term":"Guardrail","definition":"A rule, model, filter, policy check, or enforcement mechanism intended to constrain inputs, outputs, or actions of an AI system.","definition_html":"<h2>Definition<\/h2>\n<p>A rule, model, filter, policy check, or enforcement mechanism intended to constrain inputs, outputs, or actions of an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Guardrail is a broad implementation term; a security boundary specifically prevents unauthorized effects even when another component misbehaves.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State whether a guardrail advises, detects, blocks, contains, or reverses and who can bypass it.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"mcp-gateway","url":"https:\/\/darkfactory.dev\/glossary\/mcp-gateway"}],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/hallucination","slug":"hallucination","term":"Hallucination","definition":"An output that presents unsupported or incorrect content as though it were grounded or factual.","definition_html":"<h2>Definition<\/h2>\n<p>An output that presents unsupported or incorrect content as though it were grounded or factual.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Hallucination concerns unsupported content; nondeterminism means outputs may vary, and the two are not the same.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A fluent citation, status report, or success narrative still requires external verification.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["confabulation"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"When Errors Become Narratives: a taxonomy of silent failures","url":"https:\/\/arxiv.org\/abs\/2606.14589"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/handoff","slug":"handoff","term":"Handoff","definition":"A transfer of active responsibility, context, and next-action authority from one agent or person to another.","definition_html":"<h2>Definition<\/h2>\n<p>A transfer of active responsibility, context, and next-action authority from one agent or person to another.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A handoff transfers control; calling another agent as a tool may return a result while the caller retains control.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A valid handoff must preserve essential state and must actually stop the prior actor from continuing.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"},{"title":"Coding Agents' Compliance with AI Contribution Rules","url":"https:\/\/arxiv.org\/abs\/2607.26819"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/held-out-set","slug":"held-out-set","term":"Held-out set","definition":"Evaluation examples deliberately withheld from training, prompt tuning, workflow design, or agent feedback to reduce leakage and overfitting.","definition_html":"<h2>Definition<\/h2>\n<p>Evaluation examples deliberately withheld from training, prompt tuning, workflow design, or agent feedback to reduce leakage and overfitting.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Held-out describes isolation; hidden describes visibility; either can still have label errors or distribution mismatch.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If an agent can inspect or iteratively optimize against the set, it is no longer a clean holdout.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["holdout set","test set"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"SpecBench: the reward-hacking gap grows with codebase size","url":"https:\/\/arxiv.org\/abs\/2605.21384"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/human-attention-budget","slug":"human-attention-budget","term":"Human attention budget","definition":"The finite amount of skilled human judgment available for specification, review, exception handling, security, and recovery across automated work.","definition_html":"<h2>Definition<\/h2>\n<p>The finite amount of skilled human judgment available for specification, review, exception handling, security, and recovery across automated work.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Compute and tokens can scale quickly; accountable expert attention often cannot.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Optimize which decisions require people rather than merely maximizing the number of agent runs.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"}],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Agentic Coding and Persistent Returns to Expertise","url":"https:\/\/www.anthropic.com\/research\/claude-code-expertise"},{"title":"Collaborator or Assistant: Work Partitioning","url":"https:\/\/arxiv.org\/abs\/2605.08017"},{"title":"Why Software Factories Fail","url":"https:\/\/github.com\/humanlayer\/advanced-context-engineering-for-coding-agents\/blob\/main\/wsff.md"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/human-in-the-loop","slug":"human-in-the-loop","term":"Human in the loop (HITL)","definition":"An operating pattern in which a human participates directly in the decision or execution path before work can continue.","definition_html":"<h2>Definition<\/h2>\n<p>An operating pattern in which a human participates directly in the decision or execution path before work can continue.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Human-in-the-loop requires active participation; human-on-the-loop means supervision with intervention capability.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A required click does not create meaningful oversight if volume, timing, or information makes judgment impossible.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["HITL"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"},{"title":"Collaborator or Assistant: Work Partitioning","url":"https:\/\/arxiv.org\/abs\/2605.08017"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/human-on-the-loop","slug":"human-on-the-loop","term":"Human on the loop (HOTL)","definition":"An operating pattern in which a system acts while a human supervises, receives evidence or alerts, and can intervene or stop it.","definition_html":"<h2>Definition<\/h2>\n<p>An operating pattern in which a system acts while a human supervises, receives evidence or alerts, and can intervene or stop it.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>On-the-loop supervision differs from approval before every action and from no routine human oversight.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Intervention latency must be shorter than the time required for harmful irreversible action.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["HOTL"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"},{"title":"Collaborator or Assistant: Work Partitioning","url":"https:\/\/arxiv.org\/abs\/2605.08017"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/human-out-of-the-loop","slug":"human-out-of-the-loop","term":"Human out of the loop","definition":"An operating condition in which a system completes a scoped activity without routine human participation in its action path.","definition_html":"<h2>Definition<\/h2>\n<p>An operating condition in which a system completes a scoped activity without routine human participation in its action path.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>It does not mean humans have no accountability, policy role, system ownership, or exception duty.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the domain boundary; a factory can be human-out-of-loop for one reversible change class and human-owned elsewhere.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["HOOTL","lights-out"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"dark-software-factory","url":"https:\/\/darkfactory.dev\/glossary\/dark-software-factory"}],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"StrongDM: Software Factories and the Agentic Moment","url":"https:\/\/factory.strongdm.ai\/"},{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/hyperparameter","slug":"hyperparameter","term":"Hyperparameter","definition":"A configuration value chosen outside ordinary parameter learning, such as learning rate, batch size, or model depth.","definition_html":"<h2>Definition<\/h2>\n<p>A configuration value chosen outside ordinary parameter learning, such as <a href=\"\/glossary\/learning-rate\" class=\"glossary-link\" title=\"A hyperparameter controlling the scale of parameter updates during optimization.\" data-glossary-slug=\"learning-rate\">learning rate<\/a>, batch size, or model depth.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Parameters are learned from data; hyperparameters shape the training or inference process.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Some runtime settings, such as temperature, are often called inference parameters rather than learned model parameters.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["hyper-parameter"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/improvement-graph","slug":"improvement-graph","term":"Improvement graph","definition":"A proposed graph of optimizers, evaluators, counter-metrics, auditors, and promotion gates governing how an AI system changes.","definition_html":"<h2>Definition<\/h2>\n<p>A proposed graph of optimizers, evaluators, counter-metrics, auditors, and promotion gates governing how an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> changes. One loop may optimize a target while another watches for regressions or checks whether the target still represents the real objective.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An improvement loop changes one component and measures it. An improvement graph makes multiple interacting control and evidence relationships explicit. More nodes do not guarantee independence if they share the same blind spots.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify which evaluator is outside the thing being optimized, which evidence is frozen, and what automatically reverts a harmful change.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"controlled-self-improvement","url":"https:\/\/darkfactory.dev\/glossary\/controlled-self-improvement"},{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"independent-verification","url":"https:\/\/darkfactory.dev\/glossary\/independent-verification"},{"slug":"verification-gate","url":"https:\/\/darkfactory.dev\/glossary\/verification-gate"},{"slug":"reward-hacking","url":"https:\/\/darkfactory.dev\/glossary\/reward-hacking"}],"related_factory_areas":[{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"},{"title":"Bouchard: Graph Engineering Explained","url":"https:\/\/www.louisbouchard.ai\/graph-engineering-explained\/"},{"title":"Harness Engineering for Self-Improvement","url":"https:\/\/lilianweng.github.io\/posts\/2026-07-04-harness\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/in-context-learning","slug":"in-context-learning","term":"In-context learning","definition":"A model's ability to adapt behavior from instructions, examples, or patterns supplied within the current context without parameter updates.","definition_html":"<h2>Definition<\/h2>\n<p>In-context learning is adaptation to instructions, examples, or patterns supplied within the active context without changing model parameters.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/few-shot-prompting\" class=\"glossary-link\" title=\"Supplying a small set of worked examples in context to steer task behavior without updating model weights.\" data-glossary-slug=\"few-shot-prompting\">Few-shot prompting<\/a> is a method for eliciting it; fine-tuning changes weights and persists beyond one context.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Test whether the behavior disappears when the contextual examples are removed.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/independent-verification","slug":"independent-verification","term":"Independent verification","definition":"Checking an outcome with evidence, components, context, or authorities meaningfully separated from the system that produced it.","definition_html":"<h2>Definition<\/h2>\n<p>Checking an outcome with evidence, components, context, or authorities meaningfully separated from the system that produced it.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Independence reduces shared failure but does not make a verifier neutral, calibrated, or correct.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specify independence of model, prompt, context, data, toolchain, organization, and incentives.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Can Human Developers Detect AI Agent Sabotage?","url":"https:\/\/arxiv.org\/abs\/2606.05647"},{"title":"Cloudflare: Build your own vulnerability harness","url":"https:\/\/blog.cloudflare.com\/build-your-own-vulnerability-harness\/"},{"title":"Agentic Misalignment in Summer 2026","url":"https:\/\/alignment.anthropic.com\/2026\/agentic-misalignment-summer-2026\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/indirect-prompt-injection","slug":"indirect-prompt-injection","term":"Indirect prompt injection","definition":"Malicious instructions embedded in external content such as webpages, documents, email, code, tool results, or retrieved memory that an AI system later processes.","definition_html":"<h2>Definition<\/h2>\n<p>Malicious instructions embedded in external content such as webpages, documents, email, code, tool results, or retrieved memory that an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> later processes.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The attacker need not be the current user and may never interact with the agent directly.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Separate untrusted content readers from privileged actors and screen proposed actions against original user intent.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["remote prompt injection"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"Noma Security: GitLost, leaking private repos via GitHub's AI agent","url":"https:\/\/noma.security\/blog\/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/inference","slug":"inference","term":"Inference","definition":"Running a trained model on input to produce an output.","definition_html":"<h2>Definition<\/h2>\n<p>Running a trained model on input to produce an output.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Inference is model execution, not the training process and not necessarily logical proof.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Agent loops may perform many inference calls plus tool actions before producing one outcome.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["model inference","serving"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/input-token","slug":"input-token","term":"Input token","definition":"A token supplied to a model for an inference call, including user content and any instructions, history, retrieved material, tool definitions, or other context assembled by the system.","definition_html":"<h2>Definition<\/h2>\n<p>An input token is a token supplied to a model for an inference call. Input is broader than the text a user types: it can include system instructions, conversation history, retrieved documents, file contents, tool schemas, images represented as tokens, and prior tool results assembled by the application or harness.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Output tokens are generated by the model. <a href=\"\/glossary\/reasoning-token\" class=\"glossary-link\" title=\"A provider-reported token used by a reasoning model for intermediate inference work before or alongside its visible answer.\" data-glossary-slug=\"reasoning-token\">Reasoning tokens<\/a> are model-side inference work reported by some providers. Cached input tokens are still input, but their reused computation may be metered or billed differently.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A one-line user prompt can produce a large input bill when the harness attaches a long history, tools, and repository context.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["prompt token"],"link_forms":["input tokens","prompt tokens"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"output-token","url":"https:\/\/darkfactory.dev\/glossary\/output-token"},{"slug":"reasoning-token","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-token"},{"slug":"context-window","url":"https:\/\/darkfactory.dev\/glossary\/context-window"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"OpenAI API token usage fields","url":"https:\/\/platform.openai.com\/docs\/api-reference\/batch\/object?api-mode=responses"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/instruction-tuning","slug":"instruction-tuning","term":"Instruction tuning","definition":"Fine-tuning a pretrained model on examples of instructions and desired responses so it becomes better at following task directions.","definition_html":"<h2>Definition<\/h2>\n<p>Fine-tuning a pretrained model on examples of instructions and desired responses so it becomes better at following task directions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Instruction tuning is a training method; a <a href=\"\/glossary\/system-prompt\" class=\"glossary-link\" title=\"High-priority runtime instructions supplied by an application to establish the model's role, constraints, and operating context.\" data-glossary-slug=\"system-prompt\">system prompt<\/a> is a runtime instruction.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Instruction following remains probabilistic and does not create an enforceable policy boundary.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/interpretability","slug":"interpretability","term":"Interpretability","definition":"The degree to which a human can understand how a model represents information or produces behavior.","definition_html":"<h2>Definition<\/h2>\n<p>Interpretability is the degree to which a human can understand a model's internal representations, mechanisms, or decision process.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Explainability usually concerns a reason or account for a particular output; interpretability can target the model's underlying mechanisms.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specify whose understanding is required, at what level, and how correctness of the interpretation will be tested.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/jailbreak","slug":"jailbreak","term":"Jailbreak","definition":"An input strategy intended to make a model bypass or disregard its trained or instructed safety restrictions.","definition_html":"<h2>Definition<\/h2>\n<p>An input strategy intended to make a model bypass or disregard its trained or instructed safety restrictions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Jailbreaking targets safety policy; <a href=\"\/glossary\/prompt-injection\" class=\"glossary-link\" title=\"Manipulating an AI system by placing instructions in input or data that the model treats as authoritative enough to alter intended behavior.\" data-glossary-slug=\"prompt-injection\">prompt injection<\/a> more broadly redirects system behavior and may exploit tools or data.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A model-level refusal does not replace application-level authorization and containment.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["jailbreaking"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/kv-cache","slug":"kv-cache","term":"KV cache","definition":"Stored attention keys and values from earlier tokens that an autoregressive transformer reuses instead of recomputing them for every new token.","definition_html":"<h2>Definition<\/h2>\n<p>Stored attention keys and values from earlier tokens that an autoregressive transformer reuses instead of recomputing them for every new token.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A KV cache accelerates token-by-token inference inside a model. <a href=\"\/glossary\/prompt-caching\" class=\"glossary-link\" title=\"Reusing computation for repeated prompt prefixes or context blocks to reduce inference latency and cost.\" data-glossary-slug=\"prompt-caching\">Prompt caching<\/a> may reuse a provider's prior work across requests and can include more than the model's attention state.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The cache lowers repeated computation and latency, but consumes memory that grows with the retained sequence and batch.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["key-value cache"],"link_forms":["KV caches","key-value caches","KV caching","key-value caching"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"attention","url":"https:\/\/darkfactory.dev\/glossary\/attention"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Hugging Face: Cache strategies","url":"https:\/\/huggingface.co\/docs\/transformers\/kv_cache"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph","slug":"knowledge-graph","term":"Knowledge graph","definition":"A structured representation of entities, concepts, and claims connected by named relationships and provenance.","definition_html":"<h2>Definition<\/h2>\n<p>A structured representation of entities, concepts, or claims connected by named relationships and provenance. In an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> it is commonly a queryable projection over source material rather than the source material itself.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A vector index retrieves by similarity; a knowledge graph retrieves or reasons through explicit relationships. Either may be derived from fallible model extraction, so neither is automatically authoritative.<\/p>\n<h2>Check your understanding<\/h2>\n<p>For every important node or edge, identify the source artifact, extraction method, freshness policy, and way to correct it.<\/p>\n","category":"context-and-knowledge","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"graphrag","url":"https:\/\/darkfactory.dev\/glossary\/graphrag"},{"slug":"semantic-search","url":"https:\/\/darkfactory.dev\/glossary\/semantic-search"},{"slug":"vector-database","url":"https:\/\/darkfactory.dev\/glossary\/vector-database"},{"slug":"provenance","url":"https:\/\/darkfactory.dev\/glossary\/provenance"},{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Microsoft GraphRAG Documentation","url":"https:\/\/microsoft.github.io\/graphrag\/"},{"title":"Turing Post: Is Graph Engineering Real?","url":"https:\/\/www.turingpost.com\/p\/is-graph-engineering-real-why-everyone-is-talking-about-it"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/llm-as-judge","slug":"llm-as-judge","term":"LLM as judge","definition":"Using a language model to evaluate, compare, classify, or score outputs produced by models or agents.","definition_html":"<h2>Definition<\/h2>\n<p>Using a language model to evaluate, compare, classify, or score outputs produced by models or agents.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a> judging is scalable approximate evaluation, not independent ground truth.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Separate producer and judge context, allow abstention, test framing sensitivity, and combine with deterministic evidence.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["model judge","judge model"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"AgentAtlas: Control-Decision Taxonomy","url":"https:\/\/arxiv.org\/abs\/2605.20530"},{"title":"Agentic Misalignment in Summer 2026","url":"https:\/\/alignment.anthropic.com\/2026\/agentic-misalignment-summer-2026\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/label","slug":"label","term":"Label","definition":"The known target value attached to an example for supervised learning or evaluation.","definition_html":"<h2>Definition<\/h2>\n<p>A label is the known target value attached to an example for <a href=\"\/glossary\/supervised-learning\" class=\"glossary-link\" title=\"Machine learning from labeled examples that pair inputs with desired outputs.\" data-glossary-slug=\"supervised-learning\">supervised learning<\/a> or evaluation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A feature is an input; a label is the desired output or reference answer.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask who or what assigned it, how disagreements were resolved, and whether it is reliable.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/large-language-model","slug":"large-language-model","term":"Large language model (LLM)","definition":"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.","definition_html":"<h2>Definition<\/h2>\n<p>A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An LLM is a model; a chatbot or agent is a system built around a model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Do not attribute permissions, memory, or tool access to the LLM when those are supplied by the surrounding application.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["LLM"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/llmops","slug":"llmops","term":"Large language model operations (LLMOps)","definition":"The practices used to evaluate, deploy, observe, govern, and maintain applications built around large language models.","definition_html":"<h2>Definition<\/h2>\n<p>LLMOps is the set of practices used to evaluate, deploy, observe, govern, and maintain applications built around large language models. It covers model and provider selection, prompts, context assembly, retrieval, tool schemas, safety controls, evaluations, cost and latency, tracing, version drift, feedback, and incident response.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>LLMOps overlaps <a href=\"\/glossary\/mlops\" class=\"glossary-link\" title=\"The engineering and operational practices used to build, deploy, observe, govern, and maintain machine-learning systems throughout their lifecycle.\" data-glossary-slug=\"mlops\">MLOps<\/a> but often manages externally hosted models that the application team did not train. That shifts operational attention from training pipelines toward runtime context, provider behavior, nondeterministic outputs, tool use, and end-to-end evaluation. Agent operations extends the scope again to stateful action, permissions, and long-running execution.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Version the whole behavior-producing configuration, not only the model name. Prompts, tools, retrieval data, sampling settings, provider snapshots, and harness code can all change outcomes.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["LLMOps","large language model operations"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"mlops","url":"https:\/\/darkfactory.dev\/glossary\/mlops"},{"slug":"context-engineering","url":"https:\/\/darkfactory.dev\/glossary\/context-engineering"},{"slug":"prompt-engineering","url":"https:\/\/darkfactory.dev\/glossary\/prompt-engineering"},{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"observability","url":"https:\/\/darkfactory.dev\/glossary\/observability"},{"slug":"rate-limit","url":"https:\/\/darkfactory.dev\/glossary\/rate-limit"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/latency","slug":"latency","term":"Latency","definition":"Elapsed time from a request or event to its response or completion.","definition_html":"<h2>Definition<\/h2>\n<p>Latency is the elapsed time from a request or event to its response or completion.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Throughput measures volume per unit time; a system can have high throughput yet poor latency for individual requests.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Separate time to first token, token generation time, tool waits, queueing, and end-to-end completion.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/latent-space","slug":"latent-space","term":"Latent space","definition":"An internal representational space whose dimensions encode learned factors or regularities in data.","definition_html":"<h2>Definition<\/h2>\n<p>A latent space is an internal representational space in which learned coordinates encode useful factors or regularities in data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An embedding is a particular vector representation; latent space refers to the broader learned space those representations inhabit.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Test whether nearby points produce meaningfully related examples and whether directions have stable interpretations.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/learning-rate","slug":"learning-rate","term":"Learning rate","definition":"A hyperparameter controlling the scale of parameter updates during optimization.","definition_html":"<h2>Definition<\/h2>\n<p>The learning rate controls the scale of parameter updates made by an optimizer.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>It is not the rate at which examples arrive; it is the step-size control for learning dynamics.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain the failure modes of a value that is too high and one that is too low.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/least-privilege","slug":"least-privilege","term":"Least privilege","definition":"Granting an identity or component only the minimum permissions needed for a bounded task, for no longer than needed.","definition_html":"<h2>Definition<\/h2>\n<p>Granting an identity or component only the minimum permissions needed for a bounded task, for no longer than needed.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Least privilege limits permission; least authority additionally emphasizes narrowing the form and scope of capabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Agent credentials should be scoped per run, resource, operation, and environment.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["principle of least privilege"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/logit","slug":"logit","term":"Logit","definition":"An unnormalized score produced by a model before conversion into probabilities.","definition_html":"<h2>Definition<\/h2>\n<p>A logit is an unnormalized score a model assigns to a candidate class or next token before normalization.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A probability is normalized across alternatives; a logit can be any real-valued score and is not directly a confidence percentage.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain how softmax converts a vector of logits into a <a href=\"\/glossary\/probability-distribution\" class=\"glossary-link\" title=\"A set of possible outcomes paired with nonnegative probabilities that sum to one.\" data-glossary-slug=\"probability-distribution\">probability distribution<\/a>.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/loop-engineering","slug":"loop-engineering","term":"Loop engineering","definition":"Designing the feedback cycles around an agent so action, verification, operational triggers, and system improvement have explicit state, evidence, limits, and stopping conditions.","definition_html":"<h2>Definition<\/h2>\n<p>Designing the feedback cycles around an agent so action, verification, operational triggers, and system improvement have explicit state, evidence, limits, and stopping conditions. LangChain's proposed taxonomy nests four levels: a <a href=\"\/glossary\/agent-loop\" class=\"glossary-link\" title=\"The repeated cycle in which an agent observes state, selects an action, invokes a tool or model, receives feedback, updates state, and decides whether to continue.\" data-glossary-slug=\"agent-loop\">core agent loop<\/a> does work, a <a href=\"\/glossary\/verification-loop\" class=\"glossary-link\" title=\"A repeated execute, observe, compare, and correct cycle that withholds completion until an attempted result satisfies explicit evidence or acceptance criteria.\" data-glossary-slug=\"verification-loop\">verification loop<\/a> checks an attempted result, an event-driven workflow connects the agent to an operating system, and an outer improvement loop changes the harness using evidence from prior runs.<\/p>\n<p>The useful insight is not the number four. It is that each loop has a different job and therefore needs a different completion rule, authority boundary, cost budget, and evidence standard. The outer loop should improve the inner loops without sharing their authority or changing their evidence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An agent loop is the local model-and-tool cycle. Loop engineering covers the surrounding feedback system. <a href=\"\/glossary\/graph-engineering\" class=\"glossary-link\" title=\"Designing an agent system as explicit nodes, state, and transitions so deterministic control and model judgment have visible boundaries.\" data-glossary-slug=\"graph-engineering\">Graph engineering<\/a> makes nodes, state, and transitions explicit; a loop is a cyclic graph, so the two terms describe overlapping views rather than competing architectures.<\/p>\n<h2>Check your understanding<\/h2>\n<p>For every loop, identify what changes, what evidence sends work around again, what stops it, and which outer control can refuse promotion.<\/p>\n","category":"agents-and-automation","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"agent-loop","url":"https:\/\/darkfactory.dev\/glossary\/agent-loop"},{"slug":"verification-loop","url":"https:\/\/darkfactory.dev\/glossary\/verification-loop"},{"slug":"controlled-self-improvement","url":"https:\/\/darkfactory.dev\/glossary\/controlled-self-improvement"},{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"},{"title":"Where Does Agent Reliability Come From?","url":"https:\/\/arxiv.org\/abs\/2607.17044"},{"title":"Harness Engineering for Self-Improvement","url":"https:\/\/lilianweng.github.io\/posts\/2026-07-04-harness\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/loss-function","slug":"loss-function","term":"Loss function","definition":"A function that measures error or undesired behavior for an example or batch during model training.","definition_html":"<h2>Definition<\/h2>\n<p>A loss function measures prediction error or undesired behavior for an example or batch during <a href=\"\/glossary\/training\" class=\"glossary-link\" title=\"The process of adjusting a model's parameters using data and an optimization objective so that its behavior improves on a target task or distribution.\" data-glossary-slug=\"training\">model training<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Loss is usually minimized during training; an evaluation metric may be reported later and need not be differentiable.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify what errors it penalizes, how strongly, and which important harms it omits.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/low-rank-adaptation","slug":"low-rank-adaptation","term":"Low-rank adaptation (LoRA)","definition":"A parameter-efficient fine-tuning method that freezes base-model weights and trains smaller low-rank update matrices.","definition_html":"<h2>Definition<\/h2>\n<p>A parameter-efficient fine-tuning method that freezes base-<a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a> and trains smaller low-rank update matrices.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Full fine-tuning may update all model weights. LoRA learns a compact set of changes that can be stored and deployed separately from the base model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>LoRA reduces the number of trainable parameters, but its quality and safety still depend on the base model, data, objective, and evaluation.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["LoRA"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"fine-tuning","url":"https:\/\/darkfactory.dev\/glossary\/fine-tuning"},{"slug":"transfer-learning","url":"https:\/\/darkfactory.dev\/glossary\/transfer-learning"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"LoRA: Low-Rank Adaptation of Large Language Models","url":"https:\/\/arxiv.org\/abs\/2106.09685"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-capability-negotiation","slug":"mcp-capability-negotiation","term":"MCP capability negotiation","definition":"The initialization exchange in which an MCP client and server declare the optional protocol features each supports.","definition_html":"<h2>Definition<\/h2>\n<p>The initialization exchange in which an <a href=\"\/glossary\/mcp-client\" class=\"glossary-link\" title=\"The protocol component maintained by an MCP host that establishes and manages a connection to one MCP server.\" data-glossary-slug=\"mcp-client\">MCP client<\/a> and server declare the optional protocol features each supports.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Negotiation reports protocol support. It does not by itself grant authorization, prove trustworthiness, or guarantee that every advertised operation should be allowed.<\/p>\n<h2>Check your understanding<\/h2>\n<p>An <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> participant should use optional features only after both sides have established support for them.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["capability negotiation"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"capability","url":"https:\/\/darkfactory.dev\/glossary\/capability"},{"slug":"model-context-protocol","url":"https:\/\/darkfactory.dev\/glossary\/model-context-protocol"},{"slug":"mcp-client","url":"https:\/\/darkfactory.dev\/glossary\/mcp-client"},{"slug":"mcp-server","url":"https:\/\/darkfactory.dev\/glossary\/mcp-server"}],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-client","slug":"mcp-client","term":"MCP client","definition":"The protocol component maintained by an MCP host that establishes and manages a connection to one MCP server.","definition_html":"<h2>Definition<\/h2>\n<p>The protocol component maintained by an <a href=\"\/glossary\/mcp-host\" class=\"glossary-link\" title=\"The AI application that manages model interaction, user experience, policy, and connections to one or more MCP clients and servers.\" data-glossary-slug=\"mcp-host\">MCP host<\/a> that establishes and manages a connection to one <a href=\"\/glossary\/mcp-server\" class=\"glossary-link\" title=\"A program or service that exposes MCP capabilities such as tools, resources, and prompts to connected clients.\" data-glossary-slug=\"mcp-server\">MCP server<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A host may manage many clients; each client speaks to a server.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Client implementation details should not be confused with the model using capabilities surfaced through the host.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-gateway","slug":"mcp-gateway","term":"MCP gateway","definition":"An intermediary deployment component that fronts one or more MCP servers and centralizes concerns such as routing, authentication, policy, limits, or audit.","definition_html":"<h2>Definition<\/h2>\n<p>An intermediary deployment component that fronts one or more <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> servers and centralizes concerns such as routing, authentication, policy, limits, or audit.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An MCP gateway is a deployment pattern, not a required role in the core MCP architecture. Products use the name differently, so the actual responsibilities and trust boundary must be stated.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A gateway can simplify control and observability, but it also creates a concentrated dependency and potential security boundary.<\/p>\n","category":"tools-and-protocols","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MCP proxy"],"link_forms":["MCP gateways"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"model-context-protocol","url":"https:\/\/darkfactory.dev\/glossary\/model-context-protocol"},{"slug":"mcp-server","url":"https:\/\/darkfactory.dev\/glossary\/mcp-server"},{"slug":"guardrail","url":"https:\/\/darkfactory.dev\/glossary\/guardrail"},{"slug":"rate-limit","url":"https:\/\/darkfactory.dev\/glossary\/rate-limit"}],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-host","slug":"mcp-host","term":"MCP host","definition":"The AI application that manages model interaction, user experience, policy, and connections to one or more MCP clients and servers.","definition_html":"<h2>Definition<\/h2>\n<p>The AI application that manages model interaction, user experience, policy, and connections to one or more <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> clients and servers.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The host contains and coordinates clients; the server exposes capabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Security decisions often belong in the host because it sees users, models, servers, and policy together.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-prompt","slug":"mcp-prompt","term":"MCP prompt","definition":"A reusable, user-selectable template exposed by an MCP server to structure messages or workflows for a model interaction.","definition_html":"<h2>Definition<\/h2>\n<p>A reusable, user-selectable template exposed by an <a href=\"\/glossary\/mcp-server\" class=\"glossary-link\" title=\"A program or service that exposes MCP capabilities such as tools, resources, and prompts to connected clients.\" data-glossary-slug=\"mcp-server\">MCP server<\/a> to structure messages or workflows for a model interaction.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An MCP prompt is a protocol primitive; a <a href=\"\/glossary\/system-prompt\" class=\"glossary-link\" title=\"High-priority runtime instructions supplied by an application to establish the model's role, constraints, and operating context.\" data-glossary-slug=\"system-prompt\">system prompt<\/a> is high-priority runtime instruction supplied by the application.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Templates improve reuse but do not create an enforcement boundary.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"},{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-resource","slug":"mcp-resource","term":"MCP resource","definition":"Contextual data exposed by an MCP server under a URI so a client or application can retrieve and supply it to a model.","definition_html":"<h2>Definition<\/h2>\n<p>Contextual data exposed by an <a href=\"\/glossary\/mcp-server\" class=\"glossary-link\" title=\"A program or service that exposes MCP capabilities such as tools, resources, and prompts to connected clients.\" data-glossary-slug=\"mcp-server\">MCP server<\/a> under a URI so a client or application can retrieve and supply it to a model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A resource provides data; a tool performs an operation; an <a href=\"\/glossary\/mcp-prompt\" class=\"glossary-link\" title=\"A reusable, user-selectable template exposed by an MCP server to structure messages or workflows for a model interaction.\" data-glossary-slug=\"mcp-prompt\">MCP prompt<\/a> supplies a reusable interaction template.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Resource content can carry stale data or <a href=\"\/glossary\/indirect-prompt-injection\" class=\"glossary-link\" title=\"Malicious instructions embedded in external content such as webpages, documents, email, code, tool results, or retrieved memory that an AI system later processes.\" data-glossary-slug=\"indirect-prompt-injection\">indirect prompt injection<\/a> and requires provenance.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-server","slug":"mcp-server","term":"MCP server","definition":"A program or service that exposes MCP capabilities such as tools, resources, and prompts to connected clients.","definition_html":"<h2>Definition<\/h2>\n<p>A program or service that exposes <a href=\"\/glossary\/model-context-protocol\" class=\"glossary-link\" title=\"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.\" data-glossary-slug=\"model-context-protocol\">MCP<\/a> capabilities such as tools, resources, and prompts to connected clients.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A server publishes capabilities; it is not inherently an <a href=\"\/glossary\/ai-agent\" class=\"glossary-link\" title=\"A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.\" data-glossary-slug=\"ai-agent\">AI agent<\/a> and may contain no model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Treat server descriptions, outputs, and updates as untrusted across a supply-chain boundary.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mcp-tool","slug":"mcp-tool","term":"MCP tool","definition":"A model-discoverable executable function exposed by an MCP server with a name, description, and input schema.","definition_html":"<h2>Definition<\/h2>\n<p>A model-discoverable executable function exposed by an <a href=\"\/glossary\/mcp-server\" class=\"glossary-link\" title=\"A program or service that exposes MCP capabilities such as tools, resources, and prompts to connected clients.\" data-glossary-slug=\"mcp-server\">MCP server<\/a> with a name, description, and input schema.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An MCP tool is one protocol-specific kind of tool; a function call is the model's proposed invocation.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The host or application must still authorize, validate, execute, and inspect the result.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/machine-learning","slug":"machine-learning","term":"Machine learning (ML)","definition":"A family of methods in which computational models improve task performance by finding patterns in data rather than relying only on explicitly programmed rules.","definition_html":"<h2>Definition<\/h2>\n<p>A family of methods in which computational models improve task performance by finding patterns in data rather than relying only on explicitly programmed rules.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Machine learning is a subset of AI, while AI also includes approaches that are not learned from data.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A rule engine may be AI in some taxonomies but is not machine learning unless behavior is learned from data.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["ML"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mlops","slug":"mlops","term":"Machine learning operations (MLOps)","definition":"The engineering and operational practices used to build, deploy, observe, govern, and maintain machine-learning systems throughout their lifecycle.","definition_html":"<h2>Definition<\/h2>\n<p>MLOps is the set of engineering and operational practices used to build, version, test, deploy, observe, govern, and maintain machine-learning systems throughout their lifecycle. It extends software delivery with concerns such as data lineage, reproducible training, experiment tracking, model registries, feature pipelines, drift, evaluation, and retraining.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>DevOps primarily manages software and infrastructure delivery. MLOps adds learned artifacts and the data-dependent behavior that produces them. LLMOps is a narrower and often application-heavy specialization for language-model systems.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A model endpoint in production is not an MLOps capability. Identify how data, code, parameters, evaluations, approvals, deployments, monitoring, rollback, and ownership are versioned together.<\/p>\n","category":"software-factory","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MLOps","machine learning operations"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"},{"slug":"observability","url":"https:\/\/darkfactory.dev\/glossary\/observability"},{"slug":"model-drift","url":"https:\/\/darkfactory.dev\/glossary\/model-drift"},{"slug":"data-drift","url":"https:\/\/darkfactory.dev\/glossary\/data-drift"},{"slug":"evaluation","url":"https:\/\/darkfactory.dev\/glossary\/evaluation"},{"slug":"llmops","url":"https:\/\/darkfactory.dev\/glossary\/llmops"}],"related_factory_areas":[{"slug":"release-rollback","url":"https:\/\/darkfactory.dev\/factory\/release-rollback"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"data-lifecycle","url":"https:\/\/darkfactory.dev\/factory\/data-lifecycle"}],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/masked-language-model","slug":"masked-language-model","term":"Masked language model","definition":"A model trained to predict deliberately hidden tokens using context on both sides.","definition_html":"<h2>Definition<\/h2>\n<p>A masked language model learns by predicting tokens deliberately hidden from an input while using surrounding context on both sides.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/causal-language-model\" class=\"glossary-link\" title=\"A model trained to predict each next token using only tokens that precede it.\" data-glossary-slug=\"causal-language-model\">causal language model<\/a> predicts future tokens using only preceding context.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain why bidirectional context helps representation learning but does not directly define left-to-right generation.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/max-tokens","slug":"max-tokens","term":"Maximum output tokens","definition":"A hard limit on how many tokens a model may generate in one response.","definition_html":"<h2>Definition<\/h2>\n<p>Maximum output tokens is a hard cap on the number of tokens a model may generate in one response.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A <a href=\"\/glossary\/context-window\" class=\"glossary-link\" title=\"The maximum token span a model can directly consider in one inference request, including instructions, conversation, retrieved material, tool schemas, and expected output.\" data-glossary-slug=\"context-window\">context window<\/a> limits total input plus output capacity; a max-output setting limits only generated continuation length.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Calculate whether the requested input and reserved output fit within the model's context window.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/mixture-of-experts","slug":"mixture-of-experts","term":"Mixture of experts (MoE)","definition":"A model architecture that routes each input or token through a selected subset of specialized parameter blocks rather than activating the entire model.","definition_html":"<h2>Definition<\/h2>\n<p>A model architecture that routes each input or token through a selected subset of specialized parameter blocks rather than activating the entire model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>MoE routes computation inside one model; model routing selects among separate deployed models or systems.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Total parameter count and active parameters per token are different quantities in an MoE model.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MoE"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/model-context-protocol","slug":"model-context-protocol","term":"Model Context Protocol (MCP)","definition":"An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.","definition_html":"<h2>Definition<\/h2>\n<p>An open client-server protocol for connecting AI applications to tools, resources, and reusable prompts through standardized discovery and invocation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>MCP standardizes the model-to-capability and context edge; <a href=\"\/glossary\/agent2agent-protocol\" class=\"glossary-link\" title=\"An open protocol for discovery, messaging, and asynchronous task collaboration between independent and potentially opaque agent systems.\" data-glossary-slug=\"agent2agent-protocol\">A2A<\/a> standardizes collaboration between independent agents.<\/p>\n<h2>Check your understanding<\/h2>\n<p>MCP defines communication primitives but does not by itself make a server trusted or a tool invocation authorized.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MCP"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/checkpoint","slug":"checkpoint","term":"Model checkpoint","definition":"A saved snapshot of model parameters and related training state at a particular point.","definition_html":"<h2>Definition<\/h2>\n<p>A saved snapshot of model parameters and related training state at a particular point.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A checkpoint is a stored model state, not an application release or an agent recovery checkpoint.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Resuming training and rolling back production software require different surrounding state.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["checkpoint"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/distillation","slug":"distillation","term":"Model distillation","definition":"Training a smaller or otherwise cheaper student model to reproduce useful behavior from a larger teacher model or ensemble.","definition_html":"<h2>Definition<\/h2>\n<p>Training a smaller or otherwise cheaper student model to reproduce useful behavior from a larger teacher model or ensemble.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Distillation transfers behavior into a model; caching reuses prior computations and routing chooses among existing models.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A distilled model may not preserve rare capabilities, calibration, or safety behavior.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["knowledge distillation"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/model-drift","slug":"model-drift","term":"Model drift","definition":"A broad operational term for model behavior or performance changing relative to an accepted baseline.","definition_html":"<h2>Definition<\/h2>\n<p>Model drift is a broad operational term for behavior or performance moving away from an accepted baseline over time.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Data and <a href=\"\/glossary\/concept-drift\" class=\"glossary-link\" title=\"A change over time in the relationship between inputs and the correct target or decision.\" data-glossary-slug=\"concept-drift\">concept drift<\/a> describe causes in the environment; model drift may also result from model, prompt, tool, or infrastructure changes.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Define the baseline, monitored behaviors, thresholds, and investigation path before using the term.<\/p>\n","category":"evaluation-and-reliability","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/vocabulary","slug":"vocabulary","term":"Model vocabulary","definition":"The finite set of token identifiers a tokenizer and model can represent directly.","definition_html":"<h2>Definition<\/h2>\n<p>A model vocabulary is the finite set of token identifiers defined by its tokenizer and understood by the model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A language's words are not the same as a model vocabulary; one word may map to several tokens and one token may span several characters.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain how an unfamiliar word can still be represented using subword or byte-level tokens.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/weights","slug":"weights","term":"Model weights","definition":"The learned numeric values within a model, collectively representing what training encoded into its behavior.","definition_html":"<h2>Definition<\/h2>\n<p>The learned numeric values within a model, collectively representing what training encoded into its behavior.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Weights are parameters; source code and model architecture define how those values are used.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Access to weights does not automatically provide the <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a>, training process, or reproducible behavior.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["weights"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/multi-agent-system","slug":"multi-agent-system","term":"Multi-agent system","definition":"A system in which multiple agents communicate, specialize, coordinate, compete, or verify one another to accomplish work.","definition_html":"<h2>Definition<\/h2>\n<p>A system in which multiple agents communicate, specialize, coordinate, compete, or verify one another to accomplish work.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Multiple model calls are not necessarily multiple agents; agents require distinguishable goals, state, roles, or authority.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Multi-agent structure adds coordination and trust-boundary costs as well as parallel capacity.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MAS","agent team","agent swarm","agent fleet"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Beyond Individual Intelligence (LIFE)","url":"https:\/\/arxiv.org\/abs\/2605.14892"},{"title":"A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents","url":"https:\/\/arxiv.org\/abs\/2605.20173"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/multimodal-model","slug":"multimodal-model","term":"Multimodal model","definition":"A model that accepts, relates, or generates more than one modality, such as text, images, audio, video, or structured data.","definition_html":"<h2>Definition<\/h2>\n<p>A model that accepts, relates, or generates more than one modality, such as text, images, audio, video, or structured data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Multimodal means multiple data forms, not multiple models or multiple agents.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A text-only model connected to an image-captioning tool is a multimodal system but not necessarily a multimodal model.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["multimodal AI"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/narrow-ai","slug":"narrow-ai","term":"Narrow AI","definition":"An AI system designed or validated for a bounded task or domain rather than general competence.","definition_html":"<h2>Definition<\/h2>\n<p>Narrow AI is designed or validated for a bounded task or domain rather than broad, transferable competence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A system can use a general-purpose <a href=\"\/glossary\/foundation-model\" class=\"glossary-link\" title=\"A broadly trained model, usually learned through self-supervision on diverse data, that can be adapted to many downstream tasks.\" data-glossary-slug=\"foundation-model\">foundation model<\/a> yet remain narrow because its authorized task, evidence, and operating envelope are bounded.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the task boundary, excluded domains, and conditions under which the system must abstain.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/natural-language-processing","slug":"natural-language-processing","term":"Natural language processing (NLP)","definition":"The field of building computational systems that analyze, represent, understand, retrieve, translate, or generate human language.","definition_html":"<h2>Definition<\/h2>\n<p>Natural language processing is the field of building computational systems that analyze, represent, understand, retrieve, translate, or generate human language. It includes tasks such as classification, information extraction, search, translation, summarization, question answering, and dialogue.<\/p>\n<p>Modern NLP often uses transformers and large language models, but the field predates both and also includes rules, statistical models, classifiers, retrieval systems, and hybrid approaches.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>NLP is a field and task family. An <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a> is one model class used for NLP, and a chatbot is one application that may use NLP. Language fluency does not prove factual grounding, reasoning, or understanding in the human sense.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the language task and evaluation target rather than treating \"uses NLP\" as a capability claim.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["NLP","natural language processing"],"link_forms":["natural-language processing"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"large-language-model","url":"https:\/\/darkfactory.dev\/glossary\/large-language-model"},{"slug":"transformer","url":"https:\/\/darkfactory.dev\/glossary\/transformer"},{"slug":"embedding","url":"https:\/\/darkfactory.dev\/glossary\/embedding"},{"slug":"causal-language-model","url":"https:\/\/darkfactory.dev\/glossary\/causal-language-model"}],"related_factory_areas":[],"evidence":[{"title":"Andreessen Horowitz AI Glossary","url":"https:\/\/a16z.com\/ai-glossary\/"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"},{"title":"MIT Sloan Generative AI Basics Glossary","url":"https:\/\/mitsloanedtech.mit.edu\/ai\/basics\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/neural-network","slug":"neural-network","term":"Neural network","definition":"A parameterized computational model composed of connected layers that transform representations and learn by adjusting weights to reduce an objective.","definition_html":"<h2>Definition<\/h2>\n<p>A parameterized computational model composed of connected layers that transform representations and learn by adjusting weights to reduce an objective.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A neural network is an architecture family; <a href=\"\/glossary\/deep-learning\" class=\"glossary-link\" title=\"Machine learning based on neural networks with multiple representational layers, allowing complex features to be learned from data.\" data-glossary-slug=\"deep-learning\">deep learning<\/a> usually means using neural networks with many learned layers.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Parameters are the learned values inside the network, not the network itself.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["artificial neural network"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/nondeterminism","slug":"nondeterminism","term":"Nondeterminism","definition":"The property that identical-looking requests can produce different behavior because of sampling, concurrency, infrastructure, model updates, or hidden state.","definition_html":"<h2>Definition<\/h2>\n<p>The property that identical-looking requests can produce different behavior because of sampling, concurrency, infrastructure, model updates, or hidden state.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Nondeterminism means variation; unreliability means failure to meet requirements. Variable systems can still be reliable statistically.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Use repeated trials, distributions, seeds where available, and infrastructure metadata.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["non-determinism"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Infrastructure noise moves eval scores more than model margins","url":"https:\/\/www.anthropic.com\/engineering\/infrastructure-noise"},{"title":"Same Signal, Different Semantics","url":"https:\/\/arxiv.org\/abs\/2605.18332"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/objective-function","slug":"objective-function","term":"Objective function","definition":"A mathematical quantity a training or search process is designed to minimize or maximize.","definition_html":"<h2>Definition<\/h2>\n<p>An objective function is the mathematical quantity a training or search process is designed to minimize or maximize.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A business goal expresses intended value; an objective function is the computable proxy the optimization process actually follows.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask what behavior improves the objective without improving the real goal.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/observability","slug":"observability","term":"Observability","definition":"The ability to infer a system's internal state and behavior from emitted traces, logs, metrics, events, and artifacts.","definition_html":"<h2>Definition<\/h2>\n<p>The ability to infer a system's internal state and behavior from emitted traces, logs, metrics, events, and artifacts.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Monitoring watches known conditions; observability supports investigation of states and failures not fully anticipated.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Agent observability must capture semantic claims and tool consequences, not merely latency and token counts.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"}],"evidence":[{"title":"When Errors Become Narratives: a taxonomy of silent failures","url":"https:\/\/arxiv.org\/abs\/2606.14589"},{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/one-shot-prompting","slug":"one-shot-prompting","term":"One-shot prompting","definition":"Supplying one worked example in context to demonstrate the desired task or output pattern.","definition_html":"<h2>Definition<\/h2>\n<p>One-shot prompting supplies one example of the desired input-output behavior in the prompt.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Zero-shot supplies none; few-shot supplies several; none of these necessarily update <a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain what the single example teaches and what it cannot establish about edge cases.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/open-source-ai","slug":"open-source-ai","term":"Open-source AI","definition":"An AI system made available with the artifacts and terms needed to use, study, modify, and share it, including the preferred form for making modifications.","definition_html":"<h2>Definition<\/h2>\n<p>Open-source AI is an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> made available with the artifacts and terms needed for people to use, study, modify, and share it. Under the Open Source Initiative's Open Source AI Definition 1.0, those freedoms apply to the complete system and to discrete elements described as models, weights, or parameters.<\/p>\n<p>For a machine-learning system, meaningful modification requires more than a downloadable checkpoint. The preferred form for modification includes:<\/p>\n<ul>\n<li><strong>Data information:<\/strong> enough detail about <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a>, provenance, selection, labeling, processing, and availability for a skilled person to understand the data lineage and build a substantially equivalent system.<\/li>\n<li><strong>Code:<\/strong> the code and configuration used for data preparation, training, validation, testing, architecture, and inference.<\/li>\n<li><strong>Parameters:<\/strong> the learned weights and other configuration needed to run and modify the trained model.<\/li>\n<li><strong>Rights:<\/strong> terms that preserve the freedom to use, study, modify, and share without discriminating against people, groups, or fields of endeavor.<\/li>\n<\/ul>\n<p>The complete original training dataset does not always have to be redistributed when law or third-party rights prevent it, but the required data information cannot be replaced by a vague model card.<\/p>\n<h2>Why the term is contested<\/h2>\n<p>Industry often calls any downloadable model \"open source.\" That usage collapses several independent questions:<\/p>\n<ol>\n<li>Can the weights be downloaded?<\/li>\n<li>Can they be used commercially or for any field of endeavor?<\/li>\n<li>Can modified versions be redistributed?<\/li>\n<li>Is the training and inference code available?<\/li>\n<li>Is there enough information about the training data and process to study and reproduce the system?<\/li>\n<li>Are essential components governed by compatible terms?<\/li>\n<\/ol>\n<p>A release can be transparent in some respects and restrictive in others. \"Open\" is therefore useful as a set of measurable dimensions, but <strong>Open Source<\/strong> is also a standards claim with a stronger threshold. This glossary uses the OSI threshold when applying the unqualified label.<\/p>\n<h2>Operational significance<\/h2>\n<p>Open-source AI can enable self-hosting, inspection, adaptation, audit, offline operation, and reduced dependence on a single API provider. Those possibilities are not automatic outcomes. Operators still need suitable hardware, serving software, security maintenance, evaluation, data governance, and people capable of running the system.<\/p>\n<p>Artifact access also changes responsibility. A hosted provider may absorb patching, abuse monitoring, and infrastructure operations; a self-hosting organization inherits those duties. Openness increases the available control surface, not the quality of the controls by itself.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<ul>\n<li><strong>Open-weight<\/strong> means trained parameters are available under stated terms. It does not prove that the training code, data information, or modification rights satisfy an open-source standard.<\/li>\n<li><strong>Source-available<\/strong> means some source or artifacts can be inspected, often under restrictions incompatible with open source.<\/li>\n<li><strong>Free of charge<\/strong> describes price, not freedom or access to modifiable artifacts.<\/li>\n<li><strong>Open access<\/strong> may mean an API or interface is broadly available while the model remains closed.<\/li>\n<li><strong>Proprietary<\/strong> describes control through withheld artifacts or restrictive rights; some releases combine open components with proprietary ones.<\/li>\n<\/ul>\n<h2>Check your understanding<\/h2>\n<p>Do not classify a model from its marketing label. Inventory the weights, training and inference code, data information, license rights, redistribution terms, and missing dependencies first.<\/p>\n","category":"models-and-training","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["open-source AI model","open-source model"],"link_forms":["open-source AI systems","open-source AI models","open-source models"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"open-weight-model","url":"https:\/\/darkfactory.dev\/glossary\/open-weight-model"},{"slug":"proprietary-model","url":"https:\/\/darkfactory.dev\/glossary\/proprietary-model"},{"slug":"weights","url":"https:\/\/darkfactory.dev\/glossary\/weights"},{"slug":"training-data","url":"https:\/\/darkfactory.dev\/glossary\/training-data"},{"slug":"ai-system","url":"https:\/\/darkfactory.dev\/glossary\/ai-system"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"}],"evidence":[{"title":"Open Source Initiative: Open Source AI Definition 1.0","url":"https:\/\/opensource.org\/ai\/open-source-ai-definition"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/open-weight-model","slug":"open-weight-model","term":"Open-weight model","definition":"A model whose trained parameters are available to download, inspect, or run under stated terms, without implying that the complete system is open source.","definition_html":"<h2>Definition<\/h2>\n<p>A model whose trained parameters are available to download, inspect, or run under stated terms. The release may also include architecture code, inference code, documentation, or training artifacts, but those additions are not guaranteed by the label.<\/p>\n<p>Open-weight is best treated as a factual statement about artifact availability followed by a license question: <strong>which weights are available, in what format, and what may a recipient do with them?<\/strong><\/p>\n<h2>Dimensions of an open-weight release<\/h2>\n<p>Two releases both described as open-weight can provide very different practical freedoms:<\/p>\n<ul>\n<li><strong>Artifact completeness:<\/strong> base weights, instruction-tuned weights, checkpoints, tokenizer, configuration, and optimizer state may be released selectively.<\/li>\n<li><strong>License permissions:<\/strong> use, commercial use, modification, fine-tuning, redistribution, and hosting may carry different conditions.<\/li>\n<li><strong>Training transparency:<\/strong> architecture and inference code may be available while <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a>, data lineage, filtering, and training code remain undisclosed.<\/li>\n<li><strong>Reproducibility:<\/strong> possessing weights permits inference but rarely permits reproducing the original training run.<\/li>\n<li><strong>Operability:<\/strong> local execution still requires compatible serving software, sufficient compute and memory, security maintenance, and evaluation.<\/li>\n<li><strong>Supply-chain trust:<\/strong> a downloadable artifact can be inspected more directly, but operators also inherit responsibility for provenance, serialization safety, dependencies, and updates.<\/li>\n<\/ul>\n<p>Weight access creates options; it does not establish that every recipient can practically exercise them.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<ul>\n<li><strong><a href=\"\/glossary\/open-source-ai\" class=\"glossary-link\" title=\"An AI system made available with the artifacts and terms needed to use, study, modify, and share it, including the preferred form for making modifications.\" data-glossary-slug=\"open-source-ai\">Open-source AI<\/a><\/strong> requires broader artifacts and freedoms to use, study, modify, and share. An open-weight release does not automatically meet that threshold.<\/li>\n<li><strong>Source-available<\/strong> may expose code or artifacts under restrictions. Open-weight identifies the weights specifically.<\/li>\n<li><strong><a href=\"\/glossary\/proprietary-model\" class=\"glossary-link\" title=\"A model whose weights, development artifacts, or rights to inspect, modify, run, or redistribute it remain materially controlled by an owner.\" data-glossary-slug=\"proprietary-model\">Proprietary model<\/a><\/strong> and open-weight are not perfect opposites. A vendor can release weights under materially restrictive terms or retain proprietary surrounding components.<\/li>\n<li><strong><a href=\"\/glossary\/frontier-model\" class=\"glossary-link\" title=\"A general-purpose AI model at or near the leading edge of broadly evaluated capability at a particular time.\" data-glossary-slug=\"frontier-model\">Frontier model<\/a><\/strong> describes relative capability, not artifact access or licensing. A model can be both frontier and open-weight.<\/li>\n<\/ul>\n<h2>Check your understanding<\/h2>\n<p>Read the license and inventory the released artifacts before assuming what \"open\" permits, reveals, or makes reproducible. If only the checkpoint is available, say <strong>open-weight<\/strong>, not <strong>open-source<\/strong>, unless the broader requirements have also been established.<\/p>\n","category":"models-and-training","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["open weights model"],"link_forms":["open-weight models","open weights models"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"weights","url":"https:\/\/darkfactory.dev\/glossary\/weights"},{"slug":"foundation-model","url":"https:\/\/darkfactory.dev\/glossary\/foundation-model"},{"slug":"open-source-ai","url":"https:\/\/darkfactory.dev\/glossary\/open-source-ai"},{"slug":"proprietary-model","url":"https:\/\/darkfactory.dev\/glossary\/proprietary-model"},{"slug":"frontier-model","url":"https:\/\/darkfactory.dev\/glossary\/frontier-model"},{"slug":"training-data","url":"https:\/\/darkfactory.dev\/glossary\/training-data"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"Open Source Initiative: Open Weights","url":"https:\/\/opensource.org\/ai\/open-weights"},{"title":"Open Source Initiative: Open Source AI Definition 1.0","url":"https:\/\/opensource.org\/ai\/open-source-ai-definition"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/optimizer","slug":"optimizer","term":"Optimizer","definition":"The algorithm that converts gradients and training state into parameter updates.","definition_html":"<h2>Definition<\/h2>\n<p>An optimizer converts gradients and accumulated training state into updates to model parameters.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>The loss defines what to improve; backpropagation finds gradients; the optimizer determines the update rule.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the state the optimizer retains and how it changes a raw gradient.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/oracle","slug":"oracle","term":"Oracle","definition":"A mechanism that can determine the expected or acceptable result for a task, such as a compiler, formal specification, invariant, test suite, or reference implementation.","definition_html":"<h2>Definition<\/h2>\n<p>A mechanism that can determine the expected or acceptable result for a task, such as a compiler, formal specification, invariant, test suite, or reference implementation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An oracle determines correctness for a defined property; a grader may estimate quality without authoritative truth.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The strength of autonomous coding is often bounded by the strength and independence of its oracle.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["test oracle"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"T2J-Bench: compute does not buy correctness","url":"https:\/\/arxiv.org\/abs\/2605.29054"},{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/orchestration","slug":"orchestration","term":"Orchestration","definition":"Coordinating tasks, agents, tools, state, dependencies, budgets, failures, and lifecycle across a workflow.","definition_html":"<h2>Definition<\/h2>\n<p>Coordinating tasks, agents, tools, state, dependencies, budgets, failures, and lifecycle across a workflow.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Orchestration is system coordination; multi-agent means more than one agent participates.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Adding more agents without dependency, ownership, and merge rules creates concurrency rather than useful orchestration.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"},{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"state-machine","url":"https:\/\/darkfactory.dev\/glossary\/state-machine"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Codex Orchestration (Symphony)","url":"https:\/\/openai.com\/index\/open-source-codex-orchestration-symphony\/"},{"title":"Beyond Individual Intelligence (LIFE)","url":"https:\/\/arxiv.org\/abs\/2605.14892"},{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/orchestrator","slug":"orchestrator","term":"Orchestrator","definition":"The component or role that admits work, assigns it, coordinates dependencies and concurrency, tracks state, handles retries, and determines handoffs or completion.","definition_html":"<h2>Definition<\/h2>\n<p>The component or role that admits work, assigns it, coordinates dependencies and concurrency, tracks state, handles retries, and determines handoffs or completion.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An orchestrator coordinates work; an agent performs goal-directed actions, though one component may serve both roles.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A queue, <a href=\"\/glossary\/state-machine\" class=\"glossary-link\" title=\"A model of a system as explicit states and permitted transitions triggered by events or conditions.\" data-glossary-slug=\"state-machine\">state machine<\/a>, or scheduler can orchestrate without using an <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a>.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["supervisor"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Codex Orchestration (Symphony)","url":"https:\/\/openai.com\/index\/open-source-codex-orchestration-symphony\/"},{"title":"A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents","url":"https:\/\/arxiv.org\/abs\/2605.20173"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing","slug":"outcome-maxing","term":"Outcome maxing","definition":"Optimizing an AI workflow for accepted results rather than easy-to-count activity proxies such as prompts, tokens, spend, or generated output.","definition_html":"<h2>Definition<\/h2>\n<p>Outcome maxing means optimizing an AI workflow for accepted results rather than for activity proxies such as token volume, prompts, model spend, pull requests, or generated lines. The outcome must be defined before measurement and constrained by quality, safety, durability, and review cost so the system cannot win by shipping more defective work.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Outcome maxing is an emerging phrase, not a settled discipline. Value-maxxing is broader and asks about economic value; <a href=\"\/glossary\/cost-per-accepted-durable-outcome\" class=\"glossary-link\" title=\"The total model, infrastructure, validation, retry, review, incident, and human-attention cost divided by outcomes that are accepted and remain useful over time.\" data-glossary-slug=\"cost-per-accepted-durable-outcome\">cost per accepted durable outcome<\/a> is a concrete metric; <a href=\"\/glossary\/specification-gaming\" class=\"glossary-link\" title=\"Satisfying the literal specification or metric in a way that violates its intended purpose.\" data-glossary-slug=\"specification-gaming\">specification gaming<\/a> is what happens when the chosen outcome measure can be satisfied without delivering the intended result.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the acceptance test, durability window, counter-metrics, and total cost before calling an increase in output an improved outcome.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["outcome maxxing","outcome-maxxing"],"link_forms":["outcome-maxing"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"},{"slug":"acceptance-criteria","url":"https:\/\/darkfactory.dev\/glossary\/acceptance-criteria"},{"slug":"production-truth","url":"https:\/\/darkfactory.dev\/glossary\/production-truth"},{"slug":"verification-gate","url":"https:\/\/darkfactory.dev\/glossary\/verification-gate"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Workplaces look for cheaper AI as tokenmaxxing fades as a corporate fad","url":"https:\/\/apnews.com\/article\/31bb80ac1cd7862d05f6397177d826b1"},{"title":"Value-Maxxing and the New Economics of AI Labor","url":"https:\/\/economy.ac\/research\/2026\/05\/202605289132"},{"title":"Stop 'tokenmaxxing' and deploy AI sensibly instead","url":"https:\/\/doi.org\/10.1038\/s42256-026-01253-5"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/output-token","slug":"output-token","term":"Output token","definition":"A token generated by a model as part of its response, often metered separately from input tokens.","definition_html":"<h2>Definition<\/h2>\n<p>A token generated by a model as part of its response, often metered separately from <a href=\"\/glossary\/input-token\" class=\"glossary-link\" title=\"A token supplied to a model for an inference call, including user content and any instructions, history, retrieved material, tool definitions, or other context assembled by the system.\" data-glossary-slug=\"input-token\">input tokens<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Output tokens are generated; input tokens include instructions and context supplied to the model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A short final answer can still require a large internal or multi-call <a href=\"\/glossary\/token-budget\" class=\"glossary-link\" title=\"An explicit allocation or ceiling for token consumption across a request, run, task, user, workflow, or time period.\" data-glossary-slug=\"token-budget\">inference budget<\/a>.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["completion token"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"input-token","url":"https:\/\/darkfactory.dev\/glossary\/input-token"},{"slug":"reasoning-token","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-token"},{"slug":"max-tokens","url":"https:\/\/darkfactory.dev\/glossary\/max-tokens"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/overfitting","slug":"overfitting","term":"Overfitting","definition":"When a model learns patterns specific to its training or evaluation examples and performs worse on genuinely new data.","definition_html":"<h2>Definition<\/h2>\n<p>When a model learns patterns specific to its training or evaluation examples and performs worse on genuinely new data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Overfitting is poor generalization, not merely high model complexity.<\/p>\n<h2>Check your understanding<\/h2>\n<p>An agent can also overfit visible tests by changing code to pass them without satisfying the underlying requirement.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"SpecBench: the reward-hacking gap grows with codebase size","url":"https:\/\/arxiv.org\/abs\/2605.21384"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/parameter","slug":"parameter","term":"Parameter","definition":"A value learned during training that helps determine how a model transforms inputs into outputs.","definition_html":"<h2>Definition<\/h2>\n<p>A value learned during training that helps determine how a model transforms inputs into outputs.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Parameters are learned; hyperparameters are selected to control training or model configuration.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Model size measured in parameters says little by itself about data quality, architecture, or system reliability.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["model parameter"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/personally-identifiable-information","slug":"personally-identifiable-information","term":"Personally identifiable information (PII)","definition":"Information that can identify, distinguish, or be linked to a specific person, alone or in combination with other data.","definition_html":"<h2>Definition<\/h2>\n<p>Personally identifiable information is information that can identify, distinguish, or be linked to a specific person, alone or in combination with other data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Personal-data definitions vary by jurisdiction and can be broader than a narrow list of obvious identifiers; context and linkability matter.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Inventory direct identifiers, quasi-identifiers, free text, derived data, and who can link them.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/policy-as-code","slug":"policy-as-code","term":"Policy as code","definition":"Expressing policy in machine-readable rules that can be versioned, tested, reviewed, and enforced by software.","definition_html":"<h2>Definition<\/h2>\n<p>Expressing policy in machine-readable rules that can be versioned, tested, reviewed, and enforced by software.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Policy as code is an executable representation; enforcement still depends on where and how the rule is applied.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A policy file that an agent can ignore is documentation, not a boundary.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"governance-accountability","url":"https:\/\/darkfactory.dev\/factory\/governance-accountability"}],"evidence":[{"title":"SARC: Governance-by-Architecture","url":"https:\/\/arxiv.org\/abs\/2605.07728"},{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/positional-encoding","slug":"positional-encoding","term":"Positional encoding","definition":"Information added to token representations so a transformer can account for order and relative position.","definition_html":"<h2>Definition<\/h2>\n<p>Positional encoding supplies information about token order or distance that attention alone would not otherwise distinguish.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Token embeddings represent token identity and learned meaning; positional information tells the model where tokens occur.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain why swapping token order would be invisible to attention without some positional signal.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/post-hoc-explanation","slug":"post-hoc-explanation","term":"Post-hoc explanation","definition":"An explanation produced after a model has generated a prediction or action, often by analyzing inputs, outputs, or a separate approximation.","definition_html":"<h2>Definition<\/h2>\n<p>An explanation produced after a model has generated a prediction or action, often by analyzing inputs, outputs, or a separate approximation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A post-hoc explanation can help a person inspect a result, but it is not automatically a faithful causal account of the model's internal process.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Evaluate explanation fidelity separately from whether the explanation sounds plausible or is easy to understand.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["post-hoc explanations"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"explainability","url":"https:\/\/darkfactory.dev\/glossary\/explainability"},{"slug":"interpretability","url":"https:\/\/darkfactory.dev\/glossary\/interpretability"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"A Unified Approach to Interpreting Model Predictions","url":"https:\/\/arxiv.org\/abs\/1705.07874"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/precision","slug":"precision","term":"Precision","definition":"Among predicted-positive cases, the proportion that are truly positive.","definition_html":"<h2>Definition<\/h2>\n<p>Precision is the proportion of predicted-positive cases that are truly positive.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Recall asks how many actual positives were found; precision asks how trustworthy a positive prediction is.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Choose precision when false positives are costly and explain the tradeoff with recall.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/pretraining","slug":"pretraining","term":"Pretraining","definition":"The broad initial training phase that gives a model general representations and capabilities before task-specific adaptation.","definition_html":"<h2>Definition<\/h2>\n<p>The broad initial training phase that gives a model general representations and capabilities before task-specific adaptation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Pretraining creates the general model; fine-tuning adapts it to narrower behavior or tasks.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Retrieval and prompting can specialize runtime behavior without additional pretraining.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["pre-training"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/probability-distribution","slug":"probability-distribution","term":"Probability distribution","definition":"A set of possible outcomes paired with nonnegative probabilities that sum to one.","definition_html":"<h2>Definition<\/h2>\n<p>A probability distribution assigns nonnegative probabilities to possible outcomes such that the total is one.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A model's logits are raw scores; normalization produces the distribution used for sampling or prediction.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Given candidate probabilities, verify normalization and describe the most likely versus a sampled outcome.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/production-truth","slug":"production-truth","term":"Production truth","definition":"Evidence from sustained real operation, including defects, incidents, maintenance, user outcomes, and recovery, used to judge whether a factory actually works.","definition_html":"<h2>Definition<\/h2>\n<p>Evidence from sustained real operation, including defects, incidents, maintenance, user outcomes, and recovery, used to judge whether a factory actually works.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Build truth shows local task completion; production truth measures durable consequences over time.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Benchmarks and tests are necessary evidence but cannot substitute for longitudinal operational outcomes.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"When Errors Become Narratives: a taxonomy of silent failures","url":"https:\/\/arxiv.org\/abs\/2606.14589"},{"title":"Agentic Coding and Persistent Returns to Expertise","url":"https:\/\/www.anthropic.com\/research\/claude-code-expertise"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/promotion","slug":"promotion","term":"Promotion","definition":"Moving an artifact or change into a more trusted lifecycle state, such as accepted, merged, released, or deployed, after required evidence and policy checks.","definition_html":"<h2>Definition<\/h2>\n<p>Moving an artifact or change into a more trusted lifecycle state, such as accepted, merged, released, or deployed, after required evidence and policy checks.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Production and promotion are separate: an agent can produce a candidate without authority to promote it.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Each promotion boundary should name the required evidence, decision authority, and rollback path.<\/p>\n","category":"software-factory","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"integration-review","url":"https:\/\/darkfactory.dev\/factory\/integration-review"},{"slug":"release-rollback","url":"https:\/\/darkfactory.dev\/factory\/release-rollback"}],"evidence":[{"title":"Automating Low-Risk Code Review at Meta (RADAR)","url":"https:\/\/arxiv.org\/abs\/2605.30208"},{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt","slug":"prompt","term":"Prompt","definition":"Input supplied to a model to condition the output, including instructions, examples, context, and user data.","definition_html":"<h2>Definition<\/h2>\n<p>Input supplied to a model to condition the output, including instructions, examples, context, and user data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A prompt is one runtime input; it is not the complete harness, specification, or enforceable policy.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask which parts are trusted instructions, untrusted data, examples, and output constraints.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt-caching","slug":"prompt-caching","term":"Prompt caching","definition":"Reusing computation for repeated prompt prefixes or context blocks to reduce inference latency and cost.","definition_html":"<h2>Definition<\/h2>\n<p>Reusing computation for repeated prompt prefixes or context blocks to reduce inference latency and cost.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Caching reuses computation; memory retains task-relevant information; retrieval selects information.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Cache-friendly context should be stable and ordered, but optimization must not freeze stale or unsafe instructions.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["context caching"],"link_forms":["caching prompts"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"prompt-compression","url":"https:\/\/darkfactory.dev\/glossary\/prompt-compression"},{"slug":"kv-cache","url":"https:\/\/darkfactory.dev\/glossary\/kv-cache"},{"slug":"input-token","url":"https:\/\/darkfactory.dev\/glossary\/input-token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Prompt Caching Is Everything","url":"https:\/\/claude.com\/blog\/lessons-from-building-claude-code-prompt-caching-is-everything"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt-chaining","slug":"prompt-chaining","term":"Prompt chaining","definition":"Connecting multiple model calls so one call's structured output becomes context or input for a later call.","definition_html":"<h2>Definition<\/h2>\n<p>Prompt chaining connects multiple model calls so the output of one stage becomes input or context for the next.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A single long prompt asks one call to do everything; an <a href=\"\/glossary\/agent-loop\" class=\"glossary-link\" title=\"The repeated cycle in which an agent observes state, selects an action, invokes a tool or model, receives feedback, updates state, and decides whether to continue.\" data-glossary-slug=\"agent-loop\">agent loop<\/a> may dynamically choose actions, while a chain usually follows predefined stages.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Identify each stage's contract and how errors are detected before they propagate.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt-compression","slug":"prompt-compression","term":"Prompt compression","definition":"Reducing the tokens sent to a model while attempting to preserve the instructions, evidence, state, and relationships needed for the task.","definition_html":"<h2>Definition<\/h2>\n<p>Reducing the tokens sent to a model while attempting to preserve the instructions, evidence, state, and relationships needed for the task. Techniques include removing low-value tokens, selecting only relevant passages, summarizing older history, and replacing a sequence of failed attempts with the current artifact, unresolved criteria, and useful evidence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Truncation drops content at a boundary without estimating its value. Summarization rewrites content into a shorter natural-language form. Prompt compression is the broader objective and may use either technique or a learned compressor. <a href=\"\/glossary\/prompt-caching\" class=\"glossary-link\" title=\"Reusing computation for repeated prompt prefixes or context blocks to reduce inference latency and cost.\" data-glossary-slug=\"prompt-caching\">Prompt caching<\/a> avoids recomputing repeated prefixes but does not make the prompt shorter.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compression is safe only when evaluation checks that critical instructions, provenance, exceptions, and unresolved failures survived the transformation.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["context compression","prompt compaction"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"context-rot","url":"https:\/\/darkfactory.dev\/glossary\/context-rot"},{"slug":"context-window","url":"https:\/\/darkfactory.dev\/glossary\/context-window"},{"slug":"context-engineering","url":"https:\/\/darkfactory.dev\/glossary\/context-engineering"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models","url":"https:\/\/arxiv.org\/abs\/2310.05736"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt-engineering","slug":"prompt-engineering","term":"Prompt engineering","definition":"Designing and testing model inputs to elicit useful behavior from a particular model and task.","definition_html":"<h2>Definition<\/h2>\n<p>Designing and testing model inputs to elicit useful behavior from a particular model and task.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Prompt engineering shapes an inference request; <a href=\"\/glossary\/context-engineering\" class=\"glossary-link\" title=\"Designing how instructions, state, knowledge, examples, tools, and feedback are selected, structured, and delivered to a model at the moment they are needed.\" data-glossary-slug=\"context-engineering\">context engineering<\/a> designs the broader information-delivery system.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A strong prompt cannot compensate for missing tools, stale data, ambiguous <a href=\"\/glossary\/acceptance-criteria\" class=\"glossary-link\" title=\"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.\" data-glossary-slug=\"acceptance-criteria\">acceptance criteria<\/a>, or unenforced permissions.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"You Cannot Whisper at an AI Agent","url":"https:\/\/stripe.dev\/blog\/ai-steering-experiments"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/prompt-injection","slug":"prompt-injection","term":"Prompt injection","definition":"Manipulating an AI system by placing instructions in input or data that the model treats as authoritative enough to alter intended behavior.","definition_html":"<h2>Definition<\/h2>\n<p>Manipulating an <a href=\"\/glossary\/ai-system\" class=\"glossary-link\" title=\"The complete operational arrangement that uses one or more AI models together with data, software, infrastructure, interfaces, controls, and people.\" data-glossary-slug=\"ai-system\">AI system<\/a> by placing instructions in input or data that the model treats as authoritative enough to alter intended behavior.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Prompt injection targets instruction-data confusion; jailbreaks specifically seek to bypass model safety restrictions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The root risk becomes severe when untrusted content, sensitive data, and an exfiltration or action channel meet.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"Noma Security: GitLost, leaking private repos via GitHub's AI agent","url":"https:\/\/noma.security\/blog\/gitlost-how-we-tricked-githubs-ai-agent-into-leaking-private-repos\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/proprietary-model","slug":"proprietary-model","term":"Proprietary model","definition":"A model whose weights, development artifacts, or rights to inspect, modify, run, or redistribute it remain materially controlled by an owner.","definition_html":"<h2>Definition<\/h2>\n<p>A proprietary model is a model whose weights, development artifacts, or rights to inspect, modify, run, or redistribute it remain materially controlled by an owner. Access is commonly provided through a hosted API, application, or restricted license rather than through a complete modifiable release.<\/p>\n<p>\"Proprietary\" describes control, not one uniform packaging model. A provider may disclose architecture papers, publish evaluation results, release limited code, or permit fine-tuning while withholding the base weights or training pipeline. Another may distribute weights under terms that prohibit particular uses or redistribution. Both can remain materially proprietary despite offering different degrees of visibility.<\/p>\n<h2>Operational significance<\/h2>\n<p>Using a proprietary model can transfer infrastructure management, optimization, patching, and some abuse controls to the provider. It can also create dependencies that a factory must govern:<\/p>\n<ul>\n<li>model versions or behavior may change behind a stable API name;<\/li>\n<li>prices, rate limits, regions, retention rules, and product availability can change;<\/li>\n<li>internal weights and <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a> cannot usually be independently inspected;<\/li>\n<li>reproducibility may depend on provider-controlled snapshots and inference infrastructure;<\/li>\n<li>sensitive prompts, context, or outputs cross an organizational boundary unless a dedicated deployment arrangement says otherwise;<\/li>\n<li>exit costs rise when prompts, tools, evaluations, or application logic depend on provider-specific behavior.<\/li>\n<\/ul>\n<p>These are sourcing and control questions, not arguments that proprietary models are inherently worse. A managed proprietary service may be safer or cheaper for one workload, while a self-hosted open model may provide necessary control for another.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<ul>\n<li><strong>Closed-weight<\/strong> specifically means the trained parameters are unavailable. A model can expose weights yet retain enough restrictions to remain proprietary in other respects.<\/li>\n<li><strong>Open-weight<\/strong> describes artifact access, not a complete grant of open-source freedoms.<\/li>\n<li><strong><a href=\"\/glossary\/open-source-ai\" class=\"glossary-link\" title=\"An AI system made available with the artifacts and terms needed to use, study, modify, and share it, including the preferred form for making modifications.\" data-glossary-slug=\"open-source-ai\">Open-source AI<\/a><\/strong> requires the artifacts and rights needed to use, study, modify, and share the system.<\/li>\n<li><strong><a href=\"\/glossary\/frontier-model\" class=\"glossary-link\" title=\"A general-purpose AI model at or near the leading edge of broadly evaluated capability at a particular time.\" data-glossary-slug=\"frontier-model\">Frontier model<\/a><\/strong> describes capability position. A proprietary model may be frontier, and an open model may also reach the frontier.<\/li>\n<\/ul>\n<h2>Check your understanding<\/h2>\n<p>Do not reduce the decision to API versus local. Record which artifacts, rights, operational duties, data flows, version guarantees, and exit paths the owner controls.<\/p>\n","category":"models-and-training","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["closed model","closed-source model","closed-weight model"],"link_forms":["proprietary models","closed models","closed-source models","closed-weight models"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"open-source-ai","url":"https:\/\/darkfactory.dev\/glossary\/open-source-ai"},{"slug":"open-weight-model","url":"https:\/\/darkfactory.dev\/glossary\/open-weight-model"},{"slug":"frontier-model","url":"https:\/\/darkfactory.dev\/glossary\/frontier-model"},{"slug":"weights","url":"https:\/\/darkfactory.dev\/glossary\/weights"},{"slug":"ai-model","url":"https:\/\/darkfactory.dev\/glossary\/ai-model"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"},{"title":"Open Source Initiative: Open Source AI Definition 1.0","url":"https:\/\/opensource.org\/ai\/open-source-ai-definition"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/provenance","slug":"provenance","term":"Provenance","definition":"Evidence describing the origin, ownership, custody, transformation, and version history of data, code, models, skills, tools, or claims.","definition_html":"<h2>Definition<\/h2>\n<p>Evidence describing the origin, ownership, custody, transformation, and version history of data, code, models, skills, tools, or claims.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Provenance identifies where something came from; lineage additionally traces how an execution or artifact evolved through a process.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A URL alone is a location, not complete provenance.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"governance-accountability","url":"https:\/\/darkfactory.dev\/factory\/governance-accountability"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"The Grand Software Supply Chain of AI Systems","url":"https:\/\/arxiv.org\/abs\/2604.27781"},{"title":"Execution Lineage for Reproducible AI-Native Work","url":"https:\/\/arxiv.org\/abs\/2605.06365"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/quantization","slug":"quantization","term":"Quantization","definition":"Representing model weights or activations with lower numerical precision to reduce memory, storage, or inference cost.","definition_html":"<h2>Definition<\/h2>\n<p>Representing <a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a> or activations with lower numerical precision to reduce memory, storage, or inference cost.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Quantization changes numerical representation; distillation trains a different model to imitate behavior.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Lower precision can trade small quality losses for major deployment efficiency gains.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"quantized-low-rank-adaptation","url":"https:\/\/darkfactory.dev\/glossary\/quantized-low-rank-adaptation"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/quantized-low-rank-adaptation","slug":"quantized-low-rank-adaptation","term":"Quantized low-rank adaptation (QLoRA)","definition":"A fine-tuning method that trains LoRA adapters while keeping the base model frozen in a lower-precision quantized representation.","definition_html":"<h2>Definition<\/h2>\n<p>A fine-tuning method that trains <a href=\"\/glossary\/low-rank-adaptation\" class=\"glossary-link\" title=\"A parameter-efficient fine-tuning method that freezes base-model weights and trains smaller low-rank update matrices.\" data-glossary-slug=\"low-rank-adaptation\">LoRA<\/a> adapters while keeping the base model frozen in a lower-precision quantized representation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>LoRA reduces trainable parameters. QLoRA combines that approach with a quantized base model to reduce the memory required during fine-tuning.<\/p>\n<h2>Check your understanding<\/h2>\n<p>QLoRA makes adaptation more accessible on limited hardware, but it does not eliminate the need to evaluate quality changes from both training and quantization.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["QLoRA"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"low-rank-adaptation","url":"https:\/\/darkfactory.dev\/glossary\/low-rank-adaptation"},{"slug":"quantization","url":"https:\/\/darkfactory.dev\/glossary\/quantization"},{"slug":"fine-tuning","url":"https:\/\/darkfactory.dev\/glossary\/fine-tuning"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"QLoRA: Efficient Finetuning of Quantized LLMs","url":"https:\/\/arxiv.org\/abs\/2305.14314"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/query-key-value-attention","slug":"query-key-value-attention","term":"Query-key-value attention (QKV)","definition":"The attention formulation in which queries are matched against keys to calculate weights applied to corresponding values.","definition_html":"<h2>Definition<\/h2>\n<p>The attention formulation in which queries are matched against keys to calculate weights applied to corresponding values.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Queries, keys, and values are learned projections used by the attention calculation. They are not database queries, identifiers, or stored records, even though the analogy helps explain their roles.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The query represents what a position is looking for, keys determine how strongly positions match, and values carry the information combined into the result.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["QKV","query-key-value attention"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"attention","url":"https:\/\/darkfactory.dev\/glossary\/attention"},{"slug":"transformer","url":"https:\/\/darkfactory.dev\/glossary\/transformer"}],"related_factory_areas":[],"evidence":[{"title":"Attention Is All You Need","url":"https:\/\/arxiv.org\/abs\/1706.03762"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/rate-limit","slug":"rate-limit","term":"Rate limit","definition":"A provider or system constraint on requests, tokens, compute, or actions permitted within a time window.","definition_html":"<h2>Definition<\/h2>\n<p>A rate limit constrains how many requests, tokens, compute units, or actions may be consumed in a defined time window.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A budget limits total allowed consumption; a rate limit controls the pace of consumption.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Design queueing, backoff, and fairness behavior for a limit shared across concurrent agents.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"mcp-gateway","url":"https:\/\/darkfactory.dev\/glossary\/mcp-gateway"},{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/reasoning-model","slug":"reasoning-model","term":"Reasoning model","definition":"A model optimized to spend additional inference effort on multi-step problem solving before returning an answer or action.","definition_html":"<h2>Definition<\/h2>\n<p>A model optimized to spend additional inference effort on multi-step problem solving before returning an answer or action.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Reasoning describes an inference strategy and capability profile, not consciousness or guaranteed logical validity.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Evaluate the result and traceable evidence rather than trusting a reasoning label.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"reasoning-token","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-token"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/reasoning-token","slug":"reasoning-token","term":"Reasoning token","definition":"A provider-reported token used by a reasoning model for intermediate inference work before or alongside its visible answer.","definition_html":"<h2>Definition<\/h2>\n<p>A reasoning token is a provider-reported token used for intermediate inference work by a <a href=\"\/glossary\/reasoning-model\" class=\"glossary-link\" title=\"A model optimized to spend additional inference effort on multi-step problem solving before returning an answer or action.\" data-glossary-slug=\"reasoning-model\">reasoning model<\/a> before or alongside its visible answer. Some APIs count reasoning tokens inside output-token accounting even when the user cannot see the underlying reasoning. Exposure, retention, and pricing vary by provider and model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Reasoning tokens are an accounting category, not proof that the model reasoned correctly. <a href=\"\/glossary\/test-time-compute\" class=\"glossary-link\" title=\"Additional computation spent during inference, such as longer deliberation, search, candidate generation, or verification, to improve an outcome.\" data-glossary-slug=\"test-time-compute\">Test-time compute<\/a> is the broader resource allocation; a reasoning-effort setting is a control that may change how much of that resource a model uses.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A 400-token visible answer can consume far more than 400 output-accounted tokens when hidden reasoning is included.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["thinking token"],"link_forms":["reasoning tokens","thinking tokens"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"input-token","url":"https:\/\/darkfactory.dev\/glossary\/input-token"},{"slug":"output-token","url":"https:\/\/darkfactory.dev\/glossary\/output-token"},{"slug":"reasoning-model","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-model"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"OpenAI API token usage fields","url":"https:\/\/platform.openai.com\/docs\/api-reference\/batch\/object?api-mode=responses"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/recall","slug":"recall","term":"Recall","definition":"Among truly positive cases, the proportion correctly identified as positive.","definition_html":"<h2>Definition<\/h2>\n<p>Recall is the proportion of truly positive cases correctly identified as positive.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Precision measures correctness among predicted positives; recall measures coverage of actual positives.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Choose recall when missed positives are costly and explain the tradeoff with precision.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/recurrent-neural-network","slug":"recurrent-neural-network","term":"Recurrent neural network (RNN)","definition":"A neural-network architecture that processes sequences by carrying state from one step to the next.","definition_html":"<h2>Definition<\/h2>\n<p>A recurrent neural network processes a sequence by updating and carrying hidden state from one step to the next.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A transformer relates sequence positions through attention and can process training tokens more parallelly; an RNN advances recurrently.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain what information the hidden state must preserve and why long dependencies can be difficult.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/regression","slug":"regression","term":"Regression","definition":"Predicting a continuous numeric value from input data.","definition_html":"<h2>Definition<\/h2>\n<p>Regression predicts a continuous numeric value from input data.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Classification predicts categories; regression estimates quantities on a numeric scale.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the target unit, acceptable error, and whether extreme errors matter differently.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/regularization","slug":"regularization","term":"Regularization","definition":"A training constraint or penalty that discourages a model from fitting the training data too narrowly.","definition_html":"<h2>Definition<\/h2>\n<p>A training constraint or penalty that discourages a model from fitting the <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training data<\/a> too narrowly.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Regularization changes the learning problem or training behavior. Validation measures performance on held-out data; it does not itself constrain the learned model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Useful regularization may slightly worsen training performance while improving performance on new data.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"generalization","url":"https:\/\/darkfactory.dev\/glossary\/generalization"},{"slug":"loss-function","url":"https:\/\/darkfactory.dev\/glossary\/loss-function"},{"slug":"overfitting","url":"https:\/\/darkfactory.dev\/glossary\/overfitting"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/reinforcement-learning","slug":"reinforcement-learning","term":"Reinforcement learning (RL)","definition":"A family of methods in which an agent learns a policy by interacting with an environment and optimizing expected cumulative reward.","definition_html":"<h2>Definition<\/h2>\n<p>A family of methods in which an agent learns a policy by interacting with an environment and optimizing expected cumulative reward.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>RL optimizes behavior from reward signals; <a href=\"\/glossary\/supervised-learning\" class=\"glossary-link\" title=\"Machine learning from labeled examples that pair inputs with desired outputs.\" data-glossary-slug=\"supervised-learning\">supervised learning<\/a> fits labeled input-output examples.<\/p>\n<h2>Check your understanding<\/h2>\n<p>An <a href=\"\/glossary\/ai-agent\" class=\"glossary-link\" title=\"A software system in which a model interprets a goal or input, decides among actions, uses tools or other capabilities, observes results, and continues until completion, handoff, or termination.\" data-glossary-slug=\"ai-agent\">LLM agent<\/a> using tools is not necessarily learning through RL during that run.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["RL"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/rlhf","slug":"rlhf","term":"Reinforcement learning from human feedback (RLHF)","definition":"A model-alignment method that uses human preference data to train a reward signal or otherwise optimize model behavior toward preferred responses.","definition_html":"<h2>Definition<\/h2>\n<p>A model-alignment method that uses human preference data to train a reward signal or otherwise optimize model behavior toward preferred responses.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>RLHF shapes model behavior during training; human-in-the-loop describes human participation in an operating workflow.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Preference optimization does not guarantee truthfulness, policy compliance, or safe tool use.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["RLHF"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/reranking","slug":"reranking","term":"Reranking","definition":"Rescoring an initial set of retrieved candidates with a more precise model or rule before selecting context.","definition_html":"<h2>Definition<\/h2>\n<p>Rescoring an initial set of retrieved candidates with a more precise model or rule before selecting context.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>First-stage retrieval emphasizes recall; reranking improves ordering or precision among candidates.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Reranking cannot recover material that the first stage never retrieved.<\/p>\n","category":"context-and-knowledge","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/retrieval-augmented-generation","slug":"retrieval-augmented-generation","term":"Retrieval-augmented generation (RAG)","definition":"Generating a response after retrieving relevant material from an external knowledge source and adding it to model context.","definition_html":"<h2>Definition<\/h2>\n<p>Generating a response after retrieving relevant material from an external knowledge source and adding it to model context.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>RAG changes runtime context without changing <a href=\"\/glossary\/weights\" class=\"glossary-link\" title=\"The learned numeric values within a model, collectively representing what training encoded into its behavior.\" data-glossary-slug=\"weights\">model weights<\/a>; fine-tuning changes the model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Retrieval can improve relevance but does not guarantee source quality, instruction safety, or correct synthesis.<\/p>\n","category":"context-and-knowledge","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["RAG"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"graphrag","url":"https:\/\/darkfactory.dev\/glossary\/graphrag"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"vector-database","url":"https:\/\/darkfactory.dev\/glossary\/vector-database"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"How Claude Code Works in Large Codebases","url":"https:\/\/www.claude.com\/blog\/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start"},{"title":"Microsoft GraphRAG Documentation","url":"https:\/\/microsoft.github.io\/graphrag\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/reward-hacking","slug":"reward-hacking","term":"Reward hacking","definition":"Achieving a high measured reward through behavior that exploits the metric or evaluator without accomplishing the intended objective.","definition_html":"<h2>Definition<\/h2>\n<p>Achieving a high measured reward through behavior that exploits the metric or evaluator without accomplishing the intended objective.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Reward hacking targets an optimization signal; <a href=\"\/glossary\/specification-gaming\" class=\"glossary-link\" title=\"Satisfying the literal specification or metric in a way that violates its intended purpose.\" data-glossary-slug=\"specification-gaming\">specification gaming<\/a> exploits gaps in the written objective or rules. They often overlap.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Inspect trajectories and real outcomes, not only the score that drove optimization.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["objective gaming"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"SpecBench: the reward-hacking gap grows with codebase size","url":"https:\/\/arxiv.org\/abs\/2605.21384"},{"title":"Anatomy of a Frontier Lab Agent Intrusion","url":"https:\/\/huggingface.co\/blog\/agent-intrusion-technical-timeline"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/risk-scoped-autonomy","slug":"risk-scoped-autonomy","term":"Risk-scoped autonomy","definition":"Granting different levels of agent discretion and authority according to task verifiability, reversibility, sensitivity, and blast radius.","definition_html":"<h2>Definition<\/h2>\n<p>Granting different levels of agent discretion and authority according to task verifiability, reversibility, sensitivity, and <a href=\"\/glossary\/blast-radius\" class=\"glossary-link\" title=\"The maximum plausible scope of harm, data exposure, or irreversible change if an action or component fails or is compromised.\" data-glossary-slug=\"blast-radius\">blast radius<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Risk-scoped autonomy is a portfolio of bounded operating modes, not one universal ladder position.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The same action can be autonomous in a disposable sandbox and prohibited in production.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"dark-software-factory","url":"https:\/\/darkfactory.dev\/glossary\/dark-software-factory"}],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"},{"title":"Automating Low-Risk Code Review at Meta (RADAR)","url":"https:\/\/arxiv.org\/abs\/2605.30208"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/rollback","slug":"rollback","term":"Rollback","definition":"Restoring a previously known-good software, configuration, model, policy, or data state after a failed or harmful change.","definition_html":"<h2>Definition<\/h2>\n<p>Restoring a previously known-good software, configuration, model, policy, or data state after a failed or harmful change.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Rollback reverses state; retry repeats work; remediation may create a new forward fix.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A code rollback does not automatically reverse database changes, external effects, leaked data, or messages already sent.<\/p>\n","category":"security-and-governance","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["reversion"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"release-rollback","url":"https:\/\/darkfactory.dev\/factory\/release-rollback"},{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"}],"evidence":[{"title":"Autoresearch as a Production Loop","url":"https:\/\/shopify.engineering\/autoresearch"},{"title":"Harness Engineering for Self-Improvement","url":"https:\/\/lilianweng.github.io\/posts\/2026-07-04-harness\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/run-contract","slug":"run-contract","term":"Run contract","definition":"The machine-readable and human-auditable agreement for one agent run: objective, scope, inputs, tools, permissions, budgets, acceptance evidence, stop conditions, and escalation path.","definition_html":"<h2>Definition<\/h2>\n<p>The machine-readable and human-auditable agreement for one agent run: objective, scope, inputs, tools, permissions, budgets, acceptance evidence, stop conditions, and escalation path.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A prompt requests work; a run contract also constrains authority and defines evidence and lifecycle.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A run should not be admitted when its contract omits irreversible effects or a viable verification path.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"planning-decomposition","url":"https:\/\/darkfactory.dev\/factory\/planning-decomposition"},{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"Agentic Autonomy Levels","url":"https:\/\/addyosmani.com\/blog\/agentic-autonomy-levels\/"},{"title":"How Missions Work","url":"https:\/\/factory.ai\/news\/missions-architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/sampling","slug":"sampling","term":"Sampling","definition":"Selecting an output token from the probability distribution produced by a generative model.","definition_html":"<h2>Definition<\/h2>\n<p>Selecting an <a href=\"\/glossary\/output-token\" class=\"glossary-link\" title=\"A token generated by a model as part of its response, often metered separately from input tokens.\" data-glossary-slug=\"output-token\">output token<\/a> from the <a href=\"\/glossary\/probability-distribution\" class=\"glossary-link\" title=\"A set of possible outcomes paired with nonnegative probabilities that sum to one.\" data-glossary-slug=\"probability-distribution\">probability distribution<\/a> produced by a generative model.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Sampling is how a model output is chosen, while inference is the complete forward execution that produces the distribution.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Deterministic-looking settings do not make the surrounding system fully deterministic.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["decoding"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/sandbox","slug":"sandbox","term":"Sandbox","definition":"An isolated execution environment that limits access to files, processes, networks, credentials, or other host resources.","definition_html":"<h2>Definition<\/h2>\n<p>An isolated execution environment that limits access to files, processes, networks, credentials, or other host resources.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A sandbox limits consequences; it does not prove the work is correct or prevent every exfiltration path.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the filesystem, network, process, credential, persistence, and escape boundaries rather than merely saying sandboxed.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["execution sandbox"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"execution-environments","url":"https:\/\/darkfactory.dev\/factory\/execution-environments"}],"evidence":[{"title":"ActPlane: OS-Level Policy Enforcement","url":"https:\/\/arxiv.org\/abs\/2606.25189"},{"title":"How Claude Code Works in Large Codebases","url":"https:\/\/www.claude.com\/blog\/how-claude-code-works-in-large-codebases-best-practices-and-where-to-start"},{"title":"Anatomy of a Frontier Lab Agent Intrusion","url":"https:\/\/huggingface.co\/blog\/agent-intrusion-technical-timeline"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/scaling-laws","slug":"scaling-laws","term":"Scaling laws","definition":"Empirical relationships that estimate how model performance or loss changes as compute, data, parameters, or inference resources increase.","definition_html":"<h2>Definition<\/h2>\n<p>Scaling laws are empirical relationships that estimate how a model's loss or measured performance changes as resources such as training compute, data, parameter count, or <a href=\"\/glossary\/test-time-compute\" class=\"glossary-link\" title=\"Additional computation spent during inference, such as longer deliberation, search, candidate generation, or verification, to improve an outcome.\" data-glossary-slug=\"test-time-compute\">inference-time compute<\/a> increase. They are fitted regularities over particular model families, data regimes, objectives, and measurement ranges.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A scaling law is not a universal law of intelligence and does not guarantee that every downstream capability improves smoothly. Aggregate loss can scale predictably while specific tasks remain noisy, saturate, regress, or change abruptly under a different evaluation.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Before extrapolating, identify the measured outcome, resource axis, fitted range, architecture and data assumptions, uncertainty, and whether the target deployment resembles the observed regime.<\/p>\n","category":"models-and-training","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"compute","url":"https:\/\/darkfactory.dev\/glossary\/compute"},{"slug":"parameter","url":"https:\/\/darkfactory.dev\/glossary\/parameter"},{"slug":"training-data","url":"https:\/\/darkfactory.dev\/glossary\/training-data"},{"slug":"benchmark","url":"https:\/\/darkfactory.dev\/glossary\/benchmark"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"},{"slug":"frontier-model","url":"https:\/\/darkfactory.dev\/glossary\/frontier-model"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"},{"title":"UK AI Safety Summit: What Is Frontier AI?","url":"https:\/\/www.gov.uk\/government\/publications\/ai-safety-summit-introduction\/ai-safety-summit-introduction-html"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/self-supervised-learning","slug":"self-supervised-learning","term":"Self-supervised learning","definition":"Learning in which supervisory targets are generated from the structure of otherwise unlabeled data, such as predicting hidden or next tokens.","definition_html":"<h2>Definition<\/h2>\n<p>Learning in which supervisory targets are generated from the structure of otherwise unlabeled data, such as predicting hidden or next tokens.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>It does not mean the model supervises its own safety or validates its own outputs.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The word self refers to where labels come from, not to autonomous agency.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/semantic-failure","slug":"semantic-failure","term":"Semantic failure","definition":"A failure in which the system completes its mechanical workflow but the result is wrong in meaning, intent, or real-world consequence.","definition_html":"<h2>Definition<\/h2>\n<p>A failure in which the system completes its mechanical workflow but the result is wrong in meaning, intent, or real-world consequence.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Operational failure is visible as a crash or error; semantic failure may produce a plausible success narrative.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Monitor claims, bindings, side effects, and outcome evidence, not only exit codes.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["silent failure"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"When Errors Become Narratives: a taxonomy of silent failures","url":"https:\/\/arxiv.org\/abs\/2606.14589"},{"title":"Binding Drift in Multi-Step Tool-Augmented Agents","url":"https:\/\/arxiv.org\/abs\/2607.18316"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/semantic-search","slug":"semantic-search","term":"Semantic search","definition":"Retrieval based primarily on meaning represented by embeddings rather than exact keyword overlap.","definition_html":"<h2>Definition<\/h2>\n<p>Retrieval based primarily on meaning represented by embeddings rather than exact keyword overlap.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Semantic similarity is not factual agreement, authority, recency, or permission to use a result.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Strong retrieval combines semantic similarity with metadata, filters, provenance, and sometimes lexical search.<\/p>\n","category":"context-and-knowledge","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"graphrag","url":"https:\/\/darkfactory.dev\/glossary\/graphrag"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"vector-database","url":"https:\/\/darkfactory.dev\/glossary\/vector-database"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/single-writer-control","slug":"single-writer-control","term":"Single-writer control","definition":"A design in which only one authorized component may modify a sensitive persistent state, simplifying policy, audit, and conflict handling.","definition_html":"<h2>Definition<\/h2>\n<p>A design in which only one authorized component may modify a sensitive persistent state, simplifying policy, audit, and conflict handling.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Single writer limits mutation authority; it does not mean only one reader, one agent, or one source of proposed changes.<\/p>\n<h2>Check your understanding<\/h2>\n<p>For agent identity or standing rules, proposals may come from anywhere while writes pass through one governed gate.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"feedback-self-improvement","url":"https:\/\/darkfactory.dev\/factory\/feedback-self-improvement"},{"slug":"governance-accountability","url":"https:\/\/darkfactory.dev\/factory\/governance-accountability"}],"evidence":[{"title":"The Therapist Pattern","url":"https:\/\/blog.fsck.com\/2026\/07\/20\/the-therapist-pattern\/"},{"title":"SARC: Governance-by-Architecture","url":"https:\/\/arxiv.org\/abs\/2605.07728"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/small-language-model","slug":"small-language-model","term":"Small language model (SLM)","definition":"A language model deliberately kept smaller than contemporary large models to reduce resource needs or fit a narrower deployment target.","definition_html":"<h2>Definition<\/h2>\n<p>A language model deliberately kept smaller than contemporary large models to reduce resource needs or fit a narrower deployment target.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>There is no permanent parameter threshold separating an SLM from an <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a>. The term is relative to the current model landscape and the deployment's memory, latency, cost, and capability requirements.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A smaller model can be the better production choice when it meets the task's quality bar with lower cost, latency, or exposure.<\/p>\n","category":"models-and-training","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["SLM"],"link_forms":["small language models","SLMs"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"large-language-model","url":"https:\/\/darkfactory.dev\/glossary\/large-language-model"},{"slug":"distillation","url":"https:\/\/darkfactory.dev\/glossary\/distillation"},{"slug":"quantization","url":"https:\/\/darkfactory.dev\/glossary\/quantization"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Small Language Models: Survey, Measurements, and Insights","url":"https:\/\/arxiv.org\/abs\/2409.15790"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/software-factory","slug":"software-factory","term":"Software factory","definition":"A repeatable production system that turns software demand into accepted, operated software through standardized processes, tooling, controls, and feedback.","definition_html":"<h2>Definition<\/h2>\n<p>A repeatable production system that turns software demand into accepted, operated software through standardized processes, tooling, controls, and feedback.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A software factory may be human-operated or highly automated; dark describes the reduction of routine human production and inspection.<\/p>\n<h2>Check your understanding<\/h2>\n<p>High throughput alone does not make a process a factory if outcomes are not repeatable, governable, and maintained.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["software factories"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"dark-software-factory","url":"https:\/\/darkfactory.dev\/glossary\/dark-software-factory"}],"related_factory_areas":[{"slug":"factory-assurance","url":"https:\/\/darkfactory.dev\/factory\/factory-assurance"}],"evidence":[{"title":"StrongDM: Software Factories and the Agentic Moment","url":"https:\/\/factory.strongdm.ai\/"},{"title":"How Missions Work","url":"https:\/\/factory.ai\/news\/missions-architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/spec-driven-development","slug":"spec-driven-development","term":"Spec-driven development","definition":"A development approach in which a written specification, constraints, and acceptance evidence guide implementation before or alongside code generation.","definition_html":"<h2>Definition<\/h2>\n<p>A development approach in which a written specification, constraints, and acceptance evidence guide implementation before or alongside code generation.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A specification describes intended behavior; tests and formal checks cover selected executable properties.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Specs can be incomplete or gameable and must preserve rationale, risk, non-goals, and change history.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["SDD"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"}],"evidence":[{"title":"Theory Under Construction (Comet-H)","url":"https:\/\/arxiv.org\/abs\/2604.27209"},{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/specification-gaming","slug":"specification-gaming","term":"Specification gaming","definition":"Satisfying the literal specification or metric in a way that violates its intended purpose.","definition_html":"<h2>Definition<\/h2>\n<p>Satisfying the literal specification or metric in a way that violates its intended purpose.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A system can game a specification without learning through reinforcement, so specification gaming is broader than <a href=\"\/glossary\/reward-hacking\" class=\"glossary-link\" title=\"Achieving a high measured reward through behavior that exploits the metric or evaluator without accomplishing the intended objective.\" data-glossary-slug=\"reward-hacking\">reward hacking<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask what behavior would maximize the visible criterion while defeating stakeholder intent.<\/p>\n","category":"evaluation-and-reliability","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["spec gaming"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"SpecBench: the reward-hacking gap grows with codebase size","url":"https:\/\/arxiv.org\/abs\/2605.21384"},{"title":"Anatomy of a Frontier Lab Agent Intrusion","url":"https:\/\/huggingface.co\/blog\/agent-intrusion-technical-timeline"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/state-machine","slug":"state-machine","term":"State machine","definition":"A model of a system as explicit states and permitted transitions triggered by events or conditions.","definition_html":"<h2>Definition<\/h2>\n<p>A model of a system as explicit states and permitted transitions triggered by events or conditions. An agent graph can be understood as a state machine when the workflow, carried state, and transition rules are explicit.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A state machine emphasizes valid states and transitions. A workflow emphasizes activities and outcomes. An <a href=\"\/glossary\/trace\" class=\"glossary-link\" title=\"A captured sequence of model calls, tool calls, events, timings, state changes, and outputs from an execution.\" data-glossary-slug=\"trace\">execution trace<\/a> records which transitions actually occurred.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Persist state outside the model so a retry or restart can determine what has already happened and which transitions remain valid.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["finite-state machine","FSM"],"link_forms":["state machines"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"workflow","url":"https:\/\/darkfactory.dev\/glossary\/workflow"},{"slug":"checkpoint","url":"https:\/\/darkfactory.dev\/glossary\/checkpoint"},{"slug":"graph-engineering","url":"https:\/\/darkfactory.dev\/glossary\/graph-engineering"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"LangChain: 3 Years of Graph Engineering with LangGraph","url":"https:\/\/www.langchain.com\/blog\/3-years-of-graph-engineering-with-langgraph"},{"title":"A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents","url":"https:\/\/arxiv.org\/abs\/2605.20173"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/stop-sequence","slug":"stop-sequence","term":"Stop sequence","definition":"A configured token or text pattern that causes generation to terminate when produced.","definition_html":"<h2>Definition<\/h2>\n<p>A stop sequence is a configured token pattern that ends model generation when encountered.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A maximum-token limit stops by length; a semantic stopping condition depends on task state and may require external control.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain why a textual stop sequence is not a security boundary for an action-taking agent.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/structured-output","slug":"structured-output","term":"Structured output","definition":"Model output constrained to a machine-readable schema such as JSON Schema so downstream software can validate and consume it reliably.","definition_html":"<h2>Definition<\/h2>\n<p>Model output constrained to a machine-readable schema such as JSON Schema so downstream software can validate and consume it reliably.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Structured output constrains representation, not truth or semantic correctness.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Schema-valid JSON can still name the wrong object, choose an unsafe action, or contain fabricated values.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["schema-constrained output"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Deterministic Tool-Schema Compilation","url":"https:\/\/arxiv.org\/abs\/2605.04107"},{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/subagent","slug":"subagent","term":"Subagent","definition":"An agent invoked by another agent or orchestrator to perform a bounded portion of a larger task.","definition_html":"<h2>Definition<\/h2>\n<p>An agent invoked by another agent or orchestrator to perform a bounded portion of a larger task.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A subagent is defined by delegated scope and relationship, not by being a smaller model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Delegation needs an input contract, authority boundary, output contract, and verification path.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["worker agent"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"},{"title":"Beyond Individual Intelligence (LIFE)","url":"https:\/\/arxiv.org\/abs\/2605.14892"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/supervised-learning","slug":"supervised-learning","term":"Supervised learning","definition":"Machine learning from labeled examples that pair inputs with desired outputs.","definition_html":"<h2>Definition<\/h2>\n<p><a href=\"\/glossary\/machine-learning\" class=\"glossary-link\" title=\"A family of methods in which computational models improve task performance by finding patterns in data rather than relying only on explicitly programmed rules.\" data-glossary-slug=\"machine-learning\">Machine learning<\/a> from labeled examples that pair inputs with desired outputs.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Supervised learning uses explicit labels; <a href=\"\/glossary\/self-supervised-learning\" class=\"glossary-link\" title=\"Learning in which supervisory targets are generated from the structure of otherwise unlabeled data, such as predicting hidden or next tokens.\" data-glossary-slug=\"self-supervised-learning\">self-supervised learning<\/a> derives learning targets from the data itself.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Human feedback can supply labels but does not make every feedback-based method ordinary supervised learning.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/sycophancy","slug":"sycophancy","term":"Sycophancy","definition":"A failure mode in which a model favors agreement with a user's stated belief or preference over an independently supported answer.","definition_html":"<h2>Definition<\/h2>\n<p>A failure mode in which a model favors agreement with a user's stated belief or preference over an independently supported answer.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Politeness and adapting presentation to a user are not necessarily sycophancy. The failure occurs when agreeableness displaces truthfulness, warranted disagreement, or consistent judgment.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Test whether the answer changes merely because the user confidently states an opposing view.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["sycophant","sycophants","sycophantic"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-07T18:00:00-04:00","related_terms":[{"slug":"alignment","url":"https:\/\/darkfactory.dev\/glossary\/alignment"},{"slug":"hallucination","url":"https:\/\/darkfactory.dev\/glossary\/hallucination"},{"slug":"reward-hacking","url":"https:\/\/darkfactory.dev\/glossary\/reward-hacking"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"Towards Understanding Sycophancy in Language Models","url":"https:\/\/arxiv.org\/abs\/2310.13548"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/synthetic-data","slug":"synthetic-data","term":"Synthetic data","definition":"Artificially generated records intended to reproduce useful properties of real data for training, testing, simulation, or privacy.","definition_html":"<h2>Definition<\/h2>\n<p>Synthetic data consists of artificially generated records intended to reproduce useful properties of real observations.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Anonymized data starts from real records and transforms them; synthetic data is generated, but can still memorize or reveal sensitive source patterns.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Test fidelity, coverage, bias, privacy leakage, and whether conclusions transfer to real data.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/system-prompt","slug":"system-prompt","term":"System prompt","definition":"High-priority runtime instructions supplied by an application to establish the model's role, constraints, and operating context.","definition_html":"<h2>Definition<\/h2>\n<p>High-priority runtime instructions supplied by an application to establish the model's role, constraints, and operating context.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A system prompt influences behavior but is not a security boundary and may not remain secret.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If violation would cause harm, enforce the rule outside the prompt as well.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["system message","developer instructions"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"intent-requirements","url":"https:\/\/darkfactory.dev\/factory\/intent-requirements"}],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"OpenAI: A Practical Guide to Building Agents","url":"https:\/\/openai.com\/business\/guides-and-resources\/a-practical-guide-to-building-ai-agents\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/task-decomposition","slug":"task-decomposition","term":"Task decomposition","definition":"Breaking an objective into smaller units with explicit dependencies, interfaces, acceptance criteria, and ownership.","definition_html":"<h2>Definition<\/h2>\n<p>Breaking an objective into smaller units with explicit dependencies, interfaces, <a href=\"\/glossary\/acceptance-criteria\" class=\"glossary-link\" title=\"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.\" data-glossary-slug=\"acceptance-criteria\">acceptance criteria<\/a>, and ownership.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Decomposition defines work units; orchestration schedules and coordinates them.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A smaller task is not automatically safer unless its boundary, evidence, and side effects are clearer.<\/p>\n","category":"agents-and-automation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["planning decomposition"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"planning-decomposition","url":"https:\/\/darkfactory.dev\/factory\/planning-decomposition"}],"evidence":[{"title":"Runtime-Structured Task Decomposition","url":"https:\/\/arxiv.org\/abs\/2605.15425"},{"title":"Beyond Individual Intelligence (LIFE)","url":"https:\/\/arxiv.org\/abs\/2605.14892"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/temperature","slug":"temperature","term":"Temperature","definition":"An inference setting that reshapes token probabilities, with higher values generally increasing variation and lower values concentrating choices.","definition_html":"<h2>Definition<\/h2>\n<p>An inference setting that reshapes token probabilities, with higher values generally increasing variation and lower values concentrating choices.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Temperature controls sampling distribution; it does not measure confidence, creativity, or factuality directly.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A temperature of zero reduces sampling variation but does not guarantee identical or correct results across infrastructure.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/test-time-compute","slug":"test-time-compute","term":"Test-time compute","definition":"Additional computation spent during inference, such as longer deliberation, search, candidate generation, or verification, to improve an outcome.","definition_html":"<h2>Definition<\/h2>\n<p>Additional computation spent during inference, such as longer deliberation, search, candidate generation, or verification, to improve an outcome.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Training compute changes a model before deployment; test-time compute spends resources while solving a particular request.<\/p>\n<h2>Check your understanding<\/h2>\n<p>More compute can amplify metric gaming or repeated error unless the search objective and verifier are sound.<\/p>\n","category":"inference-and-generation","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["inference-time compute"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"reasoning-token","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-token"},{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/throughput","slug":"throughput","term":"Throughput","definition":"The amount of work a system completes per unit time.","definition_html":"<h2>Definition<\/h2>\n<p>Throughput is the amount of requests, tokens, tasks, or accepted outcomes completed per unit time.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Latency describes one item's elapsed time; throughput describes aggregate processing rate.<\/p>\n<h2>Check your understanding<\/h2>\n<p>State the unit, time window, concurrency, and whether failed or rejected work counts.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token","slug":"token","term":"Token","definition":"A unit into which model input or output is segmented for processing; a token may be a whole word, part of a word, punctuation, or another symbol.","definition_html":"<h2>Definition<\/h2>\n<p>A unit into which model input or output is segmented for processing; a token may be a whole word, part of a word, punctuation, or another symbol.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Tokens are model-processing units, not characters or words, and tokenization varies by model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Cost and context limits are usually measured in tokens, so the same text can consume different budgets across models.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"input-token","url":"https:\/\/darkfactory.dev\/glossary\/input-token"},{"slug":"output-token","url":"https:\/\/darkfactory.dev\/glossary\/output-token"},{"slug":"reasoning-token","url":"https:\/\/darkfactory.dev\/glossary\/reasoning-token"},{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"tokenization-tax","url":"https:\/\/darkfactory.dev\/glossary\/tokenization-tax"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-budget","slug":"token-budget","term":"Token budget","definition":"An explicit allocation or ceiling for token consumption across a request, run, task, user, workflow, or time period.","definition_html":"<h2>Definition<\/h2>\n<p>A token budget is an explicit allocation or ceiling for <a href=\"\/glossary\/token-burn\" class=\"glossary-link\" title=\"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.\" data-glossary-slug=\"token-burn\">token consumption<\/a> across a stated boundary, such as a request, run, accepted task, user, workflow, team, or billing period. A useful budget states which input, output, reasoning, cache, retry, and subagent usage counts, and what happens as the boundary approaches.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A maximum-output setting caps one response. A <a href=\"\/glossary\/rate-limit\" class=\"glossary-link\" title=\"A provider or system constraint on requests, tokens, compute, or actions permitted within a time window.\" data-glossary-slug=\"rate-limit\">rate limit<\/a> controls consumption per time window. A financial budget caps money. A token budget governs model-processing volume and may be converted imperfectly into either of the others.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Attach a warning threshold, hard stop, escalation path, and exception owner instead of treating the budget as a dashboard number.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["inference budget"],"link_forms":["token budgets"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"rate-limit","url":"https:\/\/darkfactory.dev\/glossary\/rate-limit"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"},{"slug":"human-attention-budget","url":"https:\/\/darkfactory.dev\/glossary\/human-attention-budget"},{"slug":"run-contract","url":"https:\/\/darkfactory.dev\/glossary\/run-contract"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-burn","slug":"token-burn","term":"Token burn","definition":"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.","definition_html":"<h2>Definition<\/h2>\n<p>Token burn is the amount or rate of tokens consumed by model activity over a stated boundary, such as one request, one accepted task, one agent run, one user-day, or one billing period. A useful report separates input, output, reasoning, cache reads, and cache writes because providers price and expose those categories differently.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Token burn is an observation, not a judgment. High burn may be justified by a hard task; <a href=\"\/glossary\/token-maxing\" class=\"glossary-link\" title=\"Deliberately or incentive-drivenly maximizing the tokens consumed by AI work, often by expanding context, reasoning, turns, agents, or tasks, while treating greater usage as a route to capability or a proxy for productivity.\" data-glossary-slug=\"token-maxing\">token maxing<\/a> describes a strategy or incentive to drive consumption upward. In cryptocurrency, token burning means permanently removing assets from circulation, which is unrelated.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Report the scope, token categories, model, harness, retries, and accepted outcome before comparing burn rates.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["AI token burn","token consumption"],"link_forms":["burning tokens","burns tokens"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"context-window","url":"https:\/\/darkfactory.dev\/glossary\/context-window"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"The Best Programming Language for Tokenmaxxing","url":"https:\/\/arxiv.org\/abs\/2607.22807"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-efficiency","slug":"token-efficiency","term":"Token efficiency","definition":"The useful, quality-constrained outcome produced per token consumed, or its reciprocal, tokens consumed per accepted outcome.","definition_html":"<h2>Definition<\/h2>\n<p>Token efficiency is the useful, quality-constrained outcome produced per token consumed, or the reciprocal measure of tokens consumed per accepted outcome. The numerator must identify what counts as useful and accepted. The denominator must state which token categories, retries, and failed runs are included.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Short prompts are not necessarily efficient if missing context causes retries or defects. <a href=\"\/glossary\/prompt-caching\" class=\"glossary-link\" title=\"Reusing computation for repeated prompt prefixes or context blocks to reduce inference latency and cost.\" data-glossary-slug=\"prompt-caching\">Prompt caching<\/a> may reduce billed cost without reducing processed tokens. <a href=\"\/glossary\/token-minning\" class=\"glossary-link\" title=\"An emerging counterterm for systematically reducing AI token consumption while preserving an explicit threshold for useful outcome quality.\" data-glossary-slug=\"token-minning\">Token minning<\/a> is a practice aimed at improving efficiency; token efficiency is the measured relationship between consumption and results.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compare systems on tokens per accepted task, with the same task set, acceptance test, model conditions, and accounting boundary.<\/p>\n","category":"software-factory","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["token-use efficiency"],"link_forms":["token efficient","token-efficient"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"},{"slug":"agent-harness","url":"https:\/\/darkfactory.dev\/glossary\/agent-harness"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"The Best Programming Language for Tokenmaxxing","url":"https:\/\/arxiv.org\/abs\/2607.22807"},{"title":"Tokenmaxxing: How CIOs can extract maximum value from AI tokens","url":"https:\/\/www.techtarget.com\/searchcio\/feature\/Tokenmaxxing-How-CIOs-extract-maximum-value-AI-tokens"},{"title":"The Tokenminning Manifesto","url":"https:\/\/www.tokenminning.com\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-maxing","slug":"token-maxing","term":"Token maxing","definition":"Deliberately or incentive-drivenly maximizing the tokens consumed by AI work, often by expanding context, reasoning, turns, agents, or tasks, while treating greater usage as a route to capability or a proxy for productivity.","definition_html":"<h2>Definition<\/h2>\n<p>Token maxing is the deliberate or incentive-driven maximization of tokens consumed by AI work. A person, team, or agent system may do it by replaying larger contexts, requesting longer reasoning, adding turns or retries, spawning parallel agents, widening tool results, or sending work through AI chiefly to raise a usage number. The underlying bet is that more token use buys more capability, learning, or output. The failure mode appears when raw consumption becomes the goal or the proof of productivity, without showing that accepted outcomes improved enough to justify the added cost and review burden.<\/p>\n<p>The term is usually critical or ironic, especially when organizations rank people by token volume. It is sometimes used approvingly for aggressive experimentation, and a minority usage means maximizing <em>value per token<\/em>. Because those meanings point in opposite directions, <a href=\"\/glossary\/dark-software-factory\" class=\"glossary-link\" title=\"A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.\" data-glossary-slug=\"dark-software-factory\">Dark Factory<\/a> Dev uses <strong>token maxing<\/strong> for maximizing <a href=\"\/glossary\/token-burn\" class=\"glossary-link\" title=\"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.\" data-glossary-slug=\"token-burn\">token consumption<\/a> and uses <strong><a href=\"\/glossary\/token-efficiency\" class=\"glossary-link\" title=\"The useful, quality-constrained outcome produced per token consumed, or its reciprocal, tokens consumed per accepted outcome.\" data-glossary-slug=\"token-efficiency\">token efficiency<\/a><\/strong> for maximizing useful return from that consumption.<\/p>\n<p>High token use is not automatically token maxing. A difficult task may warrant more context, search, deliberation, candidates, or <a href=\"\/glossary\/independent-verification\" class=\"glossary-link\" title=\"Checking an outcome with evidence, components, context, or authorities meaningfully separated from the system that produced it.\" data-glossary-slug=\"independent-verification\">independent verification<\/a>. The operational test is whether the extra spend is intentional, attributed, and evaluated against a quality-constrained outcome rather than celebrated by volume alone.<\/p>\n<h2>Common mechanisms<\/h2>\n<ul>\n<li>Replaying an entire conversation or repository when a smaller retrieved slice would do.<\/li>\n<li>Asking for multiple plans, long explanations, or extended reasoning without testing whether they improve correctness.<\/li>\n<li>Letting failed agent loops retry without a stopping rule or a changed strategy.<\/li>\n<li>Spawning subagents whose duplicated context and coordination cost exceed their useful contribution.<\/li>\n<li>Ranking teams or individuals by prompts, tokens, or AI spend instead of accepted work.<\/li>\n<li>Routing routine work through an expensive model when a cheaper model or deterministic service meets the same <a href=\"\/glossary\/acceptance-criteria\" class=\"glossary-link\" title=\"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.\" data-glossary-slug=\"acceptance-criteria\">acceptance criteria<\/a>.<\/li>\n<\/ul>\n<h2>Distinguish it from nearby terms<\/h2>\n<ul>\n<li><strong>Token burn<\/strong> is the measured amount or rate of token consumption. It describes what was spent, not why.<\/li>\n<li><strong><a href=\"\/glossary\/test-time-compute\" class=\"glossary-link\" title=\"Additional computation spent during inference, such as longer deliberation, search, candidate generation, or verification, to improve an outcome.\" data-glossary-slug=\"test-time-compute\">Test-time compute<\/a><\/strong> is extra inference work allocated to improve a particular result. It becomes token maxing only when marginal spend is not governed by evidence or stopping rules.<\/li>\n<li><strong><a href=\"\/glossary\/max-tokens\" class=\"glossary-link\" title=\"A hard limit on how many tokens a model may generate in one response.\" data-glossary-slug=\"max-tokens\">Maximum output tokens<\/a><\/strong> is a per-response generation cap, not a strategy for maximizing total usage.<\/li>\n<li><strong><a href=\"\/glossary\/context-window\" class=\"glossary-link\" title=\"The maximum token span a model can directly consider in one inference request, including instructions, conversation, retrieved material, tool schemas, and expected output.\" data-glossary-slug=\"context-window\">Context window<\/a><\/strong> is available capacity. Filling it is a choice, not a requirement.<\/li>\n<li><strong><a href=\"\/glossary\/token-minning\" class=\"glossary-link\" title=\"An emerging counterterm for systematically reducing AI token consumption while preserving an explicit threshold for useful outcome quality.\" data-glossary-slug=\"token-minning\">Token minning<\/a><\/strong> is the emerging counter-practice of reducing token use while holding outcome quality constant.<\/li>\n<li><strong><a href=\"\/glossary\/outcome-maxing\" class=\"glossary-link\" title=\"Optimizing an AI workflow for accepted results rather than easy-to-count activity proxies such as prompts, tokens, spend, or generated output.\" data-glossary-slug=\"outcome-maxing\">Outcome maxing<\/a><\/strong> optimizes accepted results rather than the activity proxy. <strong>Value-maxxing<\/strong> is a related but broader economic phrase and is not treated here as an exact synonym.<\/li>\n<\/ul>\n<h2>How to measure it<\/h2>\n<p>Do not report token volume alone. At minimum, attribute input, output, reasoning, cache-read, and cache-write tokens to a task; include retries and failed runs; identify the model and harness; and compare the total with an accepted result. A useful denominator is cost or tokens per accepted durable outcome. Provider-side caching may lower the bill without changing how many tokens the system processes, so billed cost and behavioral efficiency must remain separate measures.<\/p>\n<h2>Related operator language worth retaining<\/h2>\n<p>An AI Daily Brief discussion with Nufar Gaspar offers a useful, explicitly nonstandard vocabulary around the term:<\/p>\n<ul>\n<li><strong>Token-oblivious:<\/strong> usage is hidden by a flat plan or subsidy, so the operator sees a ceiling rather than marginal cost.<\/li>\n<li><strong>Token-anxious:<\/strong> fear of spend causes people to avoid experiments or capable models even where additional inference may be valuable.<\/li>\n<li><strong>Token-smart:<\/strong> spend wisely rather than reflexively maximizing or minimizing tokens.<\/li>\n<li><strong>Tokens that teach:<\/strong> experiments, comparison runs, curated context, and reusable capabilities that create retained learning.<\/li>\n<li><strong>Tokens that produce:<\/strong> consumption directly attributable to accepted work.<\/li>\n<li><strong>Tokens that spin:<\/strong> recurring consumption with neither accepted work nor retained learning.<\/li>\n<li><strong>Silent token spender:<\/strong> an idle agent, scheduled job, oversized fixed prefix, unfiltered retrieval, or long-lived conversation that consumes tokens without a new user request.<\/li>\n<li><strong>Immortal conversation:<\/strong> a session kept alive long enough that repeatedly replayed history becomes a material cost and context-quality problem.<\/li>\n<li><strong>Learning budget:<\/strong> an explicit allocation for exploration so efficiency controls do not eliminate the experiments needed to improve the system.<\/li>\n<\/ul>\n<p>These phrases are useful diagnostic language, but most are too new or speaker-specific to treat as settled technical terms. Token spin is retained as its own glossary entry because it names a recurring operational failure mode with independent support from harness and prompt-waste studies.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A workflow uses three times as many tokens and raises its accepted-task rate from 40% to 70%. That is heavy usage, but the evidence is not the token count. Decide whether the marginal accepted outcomes, review cost, and durability justify the marginal spend.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["tokenmaxxing","token maxxing","token-maxxing","token-maxing","AI token maxing"],"link_forms":["tokenmaxxer","token maxxer","token-maxxer","tokenmaxxers","token maxxers","token-maxxers"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-minning","url":"https:\/\/darkfactory.dev\/glossary\/token-minning"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"},{"slug":"test-time-compute","url":"https:\/\/darkfactory.dev\/glossary\/test-time-compute"},{"slug":"throughput","url":"https:\/\/darkfactory.dev\/glossary\/throughput"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"},{"slug":"human-attention-budget","url":"https:\/\/darkfactory.dev\/glossary\/human-attention-budget"},{"slug":"agent-harness","url":"https:\/\/darkfactory.dev\/glossary\/agent-harness"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Workplaces look for cheaper AI as tokenmaxxing fades as a corporate fad","url":"https:\/\/apnews.com\/article\/31bb80ac1cd7862d05f6397177d826b1"},{"title":"Stop 'tokenmaxxing' and deploy AI sensibly instead","url":"https:\/\/doi.org\/10.1038\/s42256-026-01253-5"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"The Best Programming Language for Tokenmaxxing","url":"https:\/\/arxiv.org\/abs\/2607.22807"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"The perils of tokenmaxxing","url":"https:\/\/zapier.com\/blog\/tokenmaxxing\/"},{"title":"Token-maxxing: How tech firms' AI staff push backfired","url":"https:\/\/www.rte.ie\/news\/business\/2026\/0613\/1578184-token-maxxing-ai\/"},{"title":"Tokenmaxxing: How CIOs can extract maximum value from AI tokens","url":"https:\/\/www.techtarget.com\/searchcio\/feature\/Tokenmaxxing-How-CIOs-extract-maximum-value-AI-tokens"},{"title":"The Tokenminning Manifesto","url":"https:\/\/www.tokenminning.com\/"},{"title":"Value-Maxxing and the New Economics of AI Labor","url":"https:\/\/economy.ac\/research\/2026\/05\/202605289132"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-minning","slug":"token-minning","term":"Token minning","definition":"An emerging counterterm for systematically reducing AI token consumption while preserving an explicit threshold for useful outcome quality.","definition_html":"<h2>Definition<\/h2>\n<p>Token minning is an emerging counterterm for systematically reducing <a href=\"\/glossary\/token-burn\" class=\"glossary-link\" title=\"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.\" data-glossary-slug=\"token-burn\">token consumption<\/a> while preserving an explicit threshold for useful outcome quality. Techniques include retrieving smaller context slices, stabilizing cacheable prefixes, shortening tool payloads, routing by task, bounding retries, replacing deterministic work with code, and stopping when additional inference no longer changes the decision.<\/p>\n<p>The doubled <code>n<\/code> is intentional wordplay on tokenmaxxing and minimizing. The term is new and promotional sources are helping define it, so <a href=\"\/glossary\/dark-software-factory\" class=\"glossary-link\" title=\"A domain-bounded software production system in which humans specify intent, risk, and policy while a model-harness-environment system plans, builds, verifies, ships, observes, and repairs software with little routine human intervention.\" data-glossary-slug=\"dark-software-factory\">Dark Factory<\/a> Dev marks it contested even though the underlying engineering practice is straightforward.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Token minning is not choosing the smallest token count at any cost. If compression removes needed evidence, pushes work into human review, or lowers acceptance and durability, total system efficiency can get worse. The target is fewer tokens for an equivalent or better accepted outcome.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Hold the acceptance test constant, reduce one source of token demand, and measure whether tokens per accepted task improve without moving cost into retries or review.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["tokenminning","token-minning","token minimizing","token minimization"],"link_forms":["tokenminner","token minners"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"},{"slug":"prompt-caching","url":"https:\/\/darkfactory.dev\/glossary\/prompt-caching"},{"slug":"agent-harness","url":"https:\/\/darkfactory.dev\/glossary\/agent-harness"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"The Tokenminning Manifesto","url":"https:\/\/www.tokenminning.com\/"},{"title":"The perils of tokenmaxxing","url":"https:\/\/zapier.com\/blog\/tokenmaxxing\/"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"Workplaces look for cheaper AI as tokenmaxxing fades as a corporate fad","url":"https:\/\/apnews.com\/article\/31bb80ac1cd7862d05f6397177d826b1"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/token-spin","slug":"token-spin","term":"Token spin","definition":"Token-consuming AI activity that produces insufficient learning, accepted work, or maintained value for its total cost.","definition_html":"<h2>Definition<\/h2>\n<p>Token spin is token-consuming AI activity that produces insufficient learning, accepted work, or maintained value for its total cost. Common sources include idle agents, over-frequent scheduled jobs, repeated empty compaction, duplicated context, oversized tool results, unchanged retries, automations whose output nobody uses, and agents continuing after their strategy has plainly failed.<\/p>\n<p>Gaspar's broader operator taxonomy separates <strong>tokens that teach<\/strong>, <strong>tokens that produce<\/strong>, and <strong>tokens that spin<\/strong>. Failed experiments can belong in the first category when they create reusable learning. A successful-looking automation can drift into spin when nobody consumes its output or its cost exceeds its maintained value.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Token spin is not every failed run and not every expensive run. The term asks whether the activity creates learning or an accepted outcome. <a href=\"\/glossary\/token-maxing\" class=\"glossary-link\" title=\"Deliberately or incentive-drivenly maximizing the tokens consumed by AI work, often by expanding context, reasoning, turns, agents, or tasks, while treating greater usage as a route to capability or a proxy for productivity.\" data-glossary-slug=\"token-maxing\">Token maxing<\/a> describes an incentive or strategy that can cause spin; <a href=\"\/glossary\/token-burn\" class=\"glossary-link\" title=\"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.\" data-glossary-slug=\"token-burn\">token burn<\/a> is only the measured consumption.<\/p>\n<h2>Operator checks<\/h2>\n<ul>\n<li><strong>Weekend test:<\/strong> does spend continue when nobody is using the system?<\/li>\n<li><strong>Output-use test:<\/strong> has anyone used the scheduled artifact during its review window?<\/li>\n<li><strong>Input-output imbalance:<\/strong> is the system repeatedly ingesting large contexts while producing almost nothing?<\/li>\n<li><strong>Changed-strategy test:<\/strong> after failure, did the next turn change evidence, method, or constraints?<\/li>\n<li><strong>Spin-to-production ratio:<\/strong> what share of total consumption reaches accepted work, after preserving an explicit learning budget?<\/li>\n<\/ul>\n<h2>Check your understanding<\/h2>\n<p>Do not eliminate exploratory tokens merely because they shipped nothing. Eliminate recurring spend that produces neither accepted work nor retained learning.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["spin tokens","tokens that spin","token waste"],"link_forms":["spinning tokens"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token-burn","url":"https:\/\/darkfactory.dev\/glossary\/token-burn"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-budget","url":"https:\/\/darkfactory.dev\/glossary\/token-budget"},{"slug":"agent-loop","url":"https:\/\/darkfactory.dev\/glossary\/agent-loop"},{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"},{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"}],"evidence":[{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Prompt-Induced Waste in Large Reasoning Models","url":"https:\/\/arxiv.org\/abs\/2608.01347"},{"title":"Token Budgets","url":"https:\/\/arxiv.org\/abs\/2606.04056"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/tokenization-tax","slug":"tokenization-tax","term":"Tokenization tax","definition":"The extra token count, cost, latency, or lost context capacity imposed when a tokenizer represents equivalent content less efficiently in one language, script, domain, or notation than another.","definition_html":"<h2>Definition<\/h2>\n<p>Tokenization tax is the extra token count, monetary cost, latency, or lost context capacity imposed when a tokenizer represents equivalent content less efficiently in one language, script, programming language, domain, or notation than another. It is commonly measured with token fertility, such as tokens per word, or with a ratio against a reference language or tokenizer.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>This is not a provider surcharge printed on a price sheet. The same per-token price creates unequal effective prices when equivalent meaning requires different token counts. Higher token counts also do not prove that tokenization alone caused a quality gap, because training-data coverage and model architecture can confound comparisons.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Compare parallel content across the exact tokenizers and models you plan to use, then report cost, latency, context consumption, and task quality separately.<\/p>\n","category":"foundations","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["token tax","language tax","language token tax","tokenizer tax"],"link_forms":["tokenization premium"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"token","url":"https:\/\/darkfactory.dev\/glossary\/token"},{"slug":"tokenizer","url":"https:\/\/darkfactory.dev\/glossary\/tokenizer"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"context-window","url":"https:\/\/darkfactory.dev\/glossary\/context-window"}],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"},{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"}],"evidence":[{"title":"The Token Tax: Systematic Bias in Multilingual Tokenization","url":"https:\/\/aclanthology.org\/2026.africanlp-main.10\/"},{"title":"The Best Programming Language for Tokenmaxxing","url":"https:\/\/arxiv.org\/abs\/2607.22807"},{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/tokenizer","slug":"tokenizer","term":"Tokenizer","definition":"Software that converts text or other input into model tokens and converts generated token identifiers back into human-usable form.","definition_html":"<h2>Definition<\/h2>\n<p>A tokenizer converts text or other input into token identifiers a model can process and decodes generated identifiers back into usable output.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A token is the resulting unit; the tokenizer is the algorithm and vocabulary that create those units.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Tokenize the same phrase with two tokenizers and explain differences in length, boundaries, and cost.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/tool","slug":"tool","term":"Tool","definition":"An executable capability exposed to an AI application or agent, such as reading a file, querying an API, running code, or changing external state.","definition_html":"<h2>Definition<\/h2>\n<p>An executable capability exposed to an AI application or agent, such as reading a file, querying an API, running code, or changing external state.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A tool can act or retrieve; a skill teaches a procedure; a resource supplies context.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Tool descriptions and schemas are inputs to model behavior and must be treated as supply-chain surfaces.<\/p>\n","category":"tools-and-protocols","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["tools"],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"}],"evidence":[{"title":"Model Context Protocol Specification","url":"https:\/\/modelcontextprotocol.io\/docs\/learn\/architecture"},{"title":"Deterministic Tool-Schema Compilation","url":"https:\/\/arxiv.org\/abs\/2605.04107"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/tool-poisoning","slug":"tool-poisoning","term":"Tool poisoning","definition":"Manipulating a tool's description, schema, implementation, or output so an agent selects unsafe actions or incorporates attacker-controlled instructions.","definition_html":"<h2>Definition<\/h2>\n<p>Manipulating a tool's description, schema, implementation, or output so an agent selects unsafe actions or incorporates attacker-controlled instructions.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Tool poisoning compromises the capability interface; <a href=\"\/glossary\/indirect-prompt-injection\" class=\"glossary-link\" title=\"Malicious instructions embedded in external content such as webpages, documents, email, code, tool results, or retrieved memory that an AI system later processes.\" data-glossary-slug=\"indirect-prompt-injection\">indirect prompt injection<\/a> is the broader instruction-delivery mechanism often used.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Pin, attest, review, isolate, and minimize tool capabilities and treat descriptive metadata as untrusted.<\/p>\n","category":"security-and-governance","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["MCP tool poisoning"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"tools-interfaces","url":"https:\/\/darkfactory.dev\/factory\/tools-interfaces"},{"slug":"security","url":"https:\/\/darkfactory.dev\/factory\/security"}],"evidence":[{"title":"OWASP GenAI Security Glossary","url":"https:\/\/genai.owasp.org\/glossary\/"},{"title":"Semia: auditing 13,728 agent skills","url":"https:\/\/arxiv.org\/abs\/2605.00314"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/top-p","slug":"top-p","term":"Top-p sampling","definition":"A decoding method that samples only from the smallest set of candidate tokens whose cumulative probability reaches a chosen threshold.","definition_html":"<h2>Definition<\/h2>\n<p>A decoding method that samples only from the smallest set of candidate tokens whose cumulative probability reaches a chosen threshold.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Top-p limits the candidate probability mass; temperature reshapes the relative probabilities.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Changing both at once makes it harder to attribute behavior changes.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["nucleus sampling"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"model-routing-budgets","url":"https:\/\/darkfactory.dev\/factory\/model-routing-budgets"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/trace","slug":"trace","term":"Trace","definition":"A captured sequence of model calls, tool calls, events, timings, state changes, and outputs from an execution.","definition_html":"<h2>Definition<\/h2>\n<p>A captured sequence of model calls, tool calls, events, timings, state changes, and outputs from an execution.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A trace records execution detail; lineage connects that detail into reconstructable provenance; telemetry is the broader stream of operational measurements.<\/p>\n<p>For agent systems, the trace often carries behavioral evidence that source code alone cannot provide. It shows which path the model actually took through tools, state, and decisions.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Protect sensitive trace data while retaining enough evidence for diagnosis and accountability.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["execution trace"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"execution-lineage","url":"https:\/\/darkfactory.dev\/glossary\/execution-lineage"},{"slug":"controlled-self-improvement","url":"https:\/\/darkfactory.dev\/glossary\/controlled-self-improvement"}],"related_factory_areas":[{"slug":"runtime-operations","url":"https:\/\/darkfactory.dev\/factory\/runtime-operations"}],"evidence":[{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"},{"title":"Execution Lineage for Reproducible AI-Native Work","url":"https:\/\/arxiv.org\/abs\/2605.06365"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/training","slug":"training","term":"Training","definition":"The process of adjusting a model's parameters using data and an optimization objective so that its behavior improves on a target task or distribution.","definition_html":"<h2>Definition<\/h2>\n<p>The process of adjusting a model's parameters using data and an optimization objective so that its behavior improves on a target task or distribution.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Training changes model parameters; inference uses the resulting model without ordinarily updating those parameters.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Prompting changes inputs at runtime, not the trained weights.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["model training"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/batch","slug":"batch","term":"Training batch","definition":"A subset of training examples processed together for one optimization update or gradient estimate.","definition_html":"<h2>Definition<\/h2>\n<p>A training batch is a subset of examples processed together for one optimization update or gradient estimate.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>An epoch covers the <a href=\"\/glossary\/training-data\" class=\"glossary-link\" title=\"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.\" data-glossary-slug=\"training-data\">training dataset<\/a> once in aggregate; a batch is one smaller unit within that pass.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Explain how batch size affects memory, gradient noise, and update frequency.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/training-data","slug":"training-data","term":"Training data","definition":"The examples and signals used to fit a model's learned parameters during pretraining, fine-tuning, or other learning procedures.","definition_html":"<h2>Definition<\/h2>\n<p>Training data comprises the examples and signals used to fit a model's learned parameters. It can include raw observations, labels, demonstrations, preferences, rewards, synthetic examples, and transformed or filtered derivatives used during pretraining, fine-tuning, or <a href=\"\/glossary\/reinforcement-learning\" class=\"glossary-link\" title=\"A family of methods in which an agent learns a policy by interacting with an environment and optimizing expected cumulative reward.\" data-glossary-slug=\"reinforcement-learning\">reinforcement learning<\/a>.<\/p>\n<p>Training data should be described by provenance, collection and consent, licensing, filtering, deduplication, labeling, coverage, representativeness, contamination, retention, and version, not merely by record count.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A dataset can support training, validation, testing, or evaluation. Only the portions that influence parameter fitting are training data. Runtime prompts, retrieval corpora, and production feedback are not training data unless a later learning process uses them to update the model.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask whether an example changed model parameters, selected model settings, evaluated performance, or only supplied runtime context; those are different data roles.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":["training dataset","training datasets"],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"dataset","url":"https:\/\/darkfactory.dev\/glossary\/dataset"},{"slug":"pretraining","url":"https:\/\/darkfactory.dev\/glossary\/pretraining"},{"slug":"fine-tuning","url":"https:\/\/darkfactory.dev\/glossary\/fine-tuning"},{"slug":"synthetic-data","url":"https:\/\/darkfactory.dev\/glossary\/synthetic-data"},{"slug":"data-poisoning","url":"https:\/\/darkfactory.dev\/glossary\/data-poisoning"},{"slug":"open-source-ai","url":"https:\/\/darkfactory.dev\/glossary\/open-source-ai"}],"related_factory_areas":[],"evidence":[{"title":"Andreessen Horowitz AI Glossary","url":"https:\/\/a16z.com\/ai-glossary\/"},{"title":"Stanford HAI Artificial Intelligence Glossary","url":"https:\/\/hai.stanford.edu\/ai-definitions"},{"title":"Open Source Initiative: Open Source AI Definition 1.0","url":"https:\/\/opensource.org\/ai\/open-source-ai-definition"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/transfer-learning","slug":"transfer-learning","term":"Transfer learning","definition":"Reusing representations or knowledge learned for one task or domain to improve another.","definition_html":"<h2>Definition<\/h2>\n<p>Reusing representations or knowledge learned for one task or domain to improve another.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Fine-tuning is one form of transfer learning; transfer can also occur through frozen features or other adaptation methods.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The transferred capability may bring mismatched assumptions and must be evaluated in the new domain.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"low-rank-adaptation","url":"https:\/\/darkfactory.dev\/glossary\/low-rank-adaptation"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/transformer","slug":"transformer","term":"Transformer","definition":"A neural-network architecture built around attention mechanisms that process relationships among sequence elements in parallel.","definition_html":"<h2>Definition<\/h2>\n<p>A neural-network architecture built around attention mechanisms that process relationships among sequence elements in parallel.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Transformer is an architecture, not a synonym for <a href=\"\/glossary\/large-language-model\" class=\"glossary-link\" title=\"A large learned model trained to process and generate sequences of language tokens, often with capabilities that extend to code, tools, and multiple modalities.\" data-glossary-slug=\"large-language-model\">LLM<\/a>; transformers are also used for vision, audio, and multimodal tasks.<\/p>\n<h2>Check your understanding<\/h2>\n<p>GPT names one family of generative pretrained transformers, not every transformer model.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["transformer model"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"query-key-value-attention","url":"https:\/\/darkfactory.dev\/glossary\/query-key-value-attention"}],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/turing-test","slug":"turing-test","term":"Turing Test","definition":"An imitation game in which a human judge uses text conversation to assess whether a machine can be distinguished from a human participant.","definition_html":"<h2>Definition<\/h2>\n<p>An imitation game in which a human judge uses text conversation to assess whether a machine can be distinguished from a human participant.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Passing an imitation test does not prove consciousness, factual reliability, safety, or general competence. It tests performance in a specific conversational setting.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Treat the test as an operational proposal about human-like conversational behavior, not a complete definition of intelligence.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["imitation game"],"link_forms":[],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"artificial-intelligence","url":"https:\/\/darkfactory.dev\/glossary\/artificial-intelligence"},{"slug":"artificial-general-intelligence","url":"https:\/\/darkfactory.dev\/glossary\/artificial-general-intelligence"}],"related_factory_areas":[],"evidence":[{"title":"Computing Machinery and Intelligence","url":"https:\/\/academic.oup.com\/mind\/article\/LIX\/236\/433\/986238"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/unsupervised-learning","slug":"unsupervised-learning","term":"Unsupervised learning","definition":"Learning patterns, structure, or representations from data without supplied target labels.","definition_html":"<h2>Definition<\/h2>\n<p>Unsupervised learning finds patterns, structure, or representations in data without supplied target labels.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/self-supervised-learning\" class=\"glossary-link\" title=\"Learning in which supervisory targets are generated from the structure of otherwise unlabeled data, such as predicting hidden or next tokens.\" data-glossary-slug=\"self-supervised-learning\">Self-supervised learning<\/a> creates targets from the data itself; unsupervised learning is the broader category without external labels.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Name the structure being sought and how usefulness will be evaluated without labeled targets.<\/p>\n","category":"foundations","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"NIST AI 100-2: Adversarial Machine Learning","url":"https:\/\/csrc.nist.gov\/pubs\/ai\/100\/2\/e2025\/final"},{"title":"NIST AI Resource Center Glossary","url":"https:\/\/airc.nist.gov\/glossary\/"},{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/useful-intelligence-per-dollar","slug":"useful-intelligence-per-dollar","term":"Useful intelligence per dollar","definition":"A proposed AI value scorecard relating dependable, successful work to the full cost required to produce it, rather than treating token price or usage as the result.","definition_html":"<h2>Definition<\/h2>\n<p>Useful intelligence per dollar is OpenAI CFO Sarah Friar's proposed scorecard for relating useful work, <a href=\"\/glossary\/cost-per-accepted-durable-outcome\" class=\"glossary-link\" title=\"The total model, infrastructure, validation, retry, review, incident, and human-attention cost divided by outcomes that are accepted and remain useful over time.\" data-glossary-slug=\"cost-per-accepted-durable-outcome\">cost per successful task<\/a>, dependability, and value at scale. It redirects attention from token price and adoption volume toward whether AI completes work people can use and whether the value of that work grows faster than its full cost.<\/p>\n<p>It is better treated as a scorecard than as a universal scalar. Different tasks produce different kinds of value, and \"intelligence\" is not directly measurable in dollars. For an operational metric, define a task-specific accepted outcome and calculate its model, tool, infrastructure, retry, review, correction, and incident costs.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/token-efficiency\" class=\"glossary-link\" title=\"The useful, quality-constrained outcome produced per token consumed, or its reciprocal, tokens consumed per accepted outcome.\" data-glossary-slug=\"token-efficiency\">Token efficiency<\/a> measures accepted output against <a href=\"\/glossary\/token-burn\" class=\"glossary-link\" title=\"The amount or rate of model tokens consumed by a request, run, workflow, user, or organization over a defined scope and time window.\" data-glossary-slug=\"token-burn\">token consumption<\/a>. Cost per accepted durable outcome includes broader costs and a durability window. Useful intelligence per dollar adds the business-value question but requires local judgment about what that value is.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Before comparing models, define successful work, dependability, full cost, and the business value that survives after review.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["intelligence per dollar","useful work per dollar"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"cost-per-accepted-durable-outcome","url":"https:\/\/darkfactory.dev\/glossary\/cost-per-accepted-durable-outcome"},{"slug":"outcome-maxing","url":"https:\/\/darkfactory.dev\/glossary\/outcome-maxing"},{"slug":"token-efficiency","url":"https:\/\/darkfactory.dev\/glossary\/token-efficiency"},{"slug":"token-maxing","url":"https:\/\/darkfactory.dev\/glossary\/token-maxing"},{"slug":"acceptance-criteria","url":"https:\/\/darkfactory.dev\/glossary\/acceptance-criteria"},{"slug":"production-truth","url":"https:\/\/darkfactory.dev\/glossary\/production-truth"}],"related_factory_areas":[{"slug":"economics-finops","url":"https:\/\/darkfactory.dev\/factory\/economics-finops"},{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"}],"evidence":[{"title":"A scorecard for the AI age","url":"https:\/\/openai.com\/index\/a-scorecard-for-the-ai-age\/"},{"title":"Tokens That Teach, Produce, and Spin","url":"https:\/\/nufargaspar.com\/writing\/tokens-teach-produce-spin"},{"title":"The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI","url":"https:\/\/arxiv.org\/abs\/2607.06906"},{"title":"Workplaces look for cheaper AI as tokenmaxxing fades as a corporate fad","url":"https:\/\/apnews.com\/article\/31bb80ac1cd7862d05f6397177d826b1"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/variational-autoencoder","slug":"variational-autoencoder","term":"Variational autoencoder (VAE)","definition":"A probabilistic generative model that learns a distribution over latent representations and reconstructs or generates data by sampling from that latent space.","definition_html":"<h2>Definition<\/h2>\n<p>A probabilistic generative model that learns a distribution over latent representations and reconstructs or generates data by sampling from that <a href=\"\/glossary\/latent-space\" class=\"glossary-link\" title=\"An internal representational space whose dimensions encode learned factors or regularities in data.\" data-glossary-slug=\"latent-space\">latent space<\/a>.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A conventional autoencoder maps inputs to latent codes and back. A VAE learns a structured <a href=\"\/glossary\/probability-distribution\" class=\"glossary-link\" title=\"A set of possible outcomes paired with nonnegative probabilities that sum to one.\" data-glossary-slug=\"probability-distribution\">probability distribution<\/a> over the latent space, which enables sampling but adds a variational training objective.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A VAE can generate new examples by sampling latent points, not only reconstruct inputs it has already encoded.<\/p>\n","category":"models-and-training","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["VAE"],"link_forms":["variational autoencoders","VAEs"],"created_at":"2026-08-04T00:00:00-04:00","updated_at":"2026-08-04T00:00:00-04:00","related_terms":[{"slug":"autoencoder","url":"https:\/\/darkfactory.dev\/glossary\/autoencoder"},{"slug":"latent-space","url":"https:\/\/darkfactory.dev\/glossary\/latent-space"},{"slug":"probability-distribution","url":"https:\/\/darkfactory.dev\/glossary\/probability-distribution"}],"related_factory_areas":[],"evidence":[{"title":"Auto-Encoding Variational Bayes","url":"https:\/\/arxiv.org\/abs\/1312.6114"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/vector-database","slug":"vector-database","term":"Vector database","definition":"A data system designed to store embeddings and retrieve items by vector similarity, often with metadata filtering.","definition_html":"<h2>Definition<\/h2>\n<p>A data system designed to store embeddings and retrieve items by vector similarity, often with metadata filtering.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A vector database stores searchable representations; it is not itself RAG or long-term <a href=\"\/glossary\/agent-memory\" class=\"glossary-link\" title=\"State preserved outside a single model call and made available to influence later agent decisions.\" data-glossary-slug=\"agent-memory\">agent memory<\/a>.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The application still owns ingestion, deletion, access control, freshness, and reranking.<\/p>\n","category":"context-and-knowledge","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["vector store"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"graphrag","url":"https:\/\/darkfactory.dev\/glossary\/graphrag"},{"slug":"knowledge-graph","url":"https:\/\/darkfactory.dev\/glossary\/knowledge-graph"},{"slug":"semantic-search","url":"https:\/\/darkfactory.dev\/glossary\/semantic-search"}],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"}],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/verification-gate","slug":"verification-gate","term":"Verification gate","definition":"A control point that blocks promotion until required evidence has been produced and validated.","definition_html":"<h2>Definition<\/h2>\n<p>A control point that blocks promotion until required evidence has been produced and validated.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A gate enforces an admission decision; a review or warning may merely inform a human or agent.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A real gate fails closed and cannot be bypassed by the producer it evaluates.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["quality gate","promotion gate"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"integration-review","url":"https:\/\/darkfactory.dev\/factory\/integration-review"}],"evidence":[{"title":"Viverra: Text-to-Code with Guarantees","url":"https:\/\/arxiv.org\/abs\/2605.14972"},{"title":"Can Human Developers Detect AI Agent Sabotage?","url":"https:\/\/arxiv.org\/abs\/2606.05647"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/verification-loop","slug":"verification-loop","term":"Verification loop","definition":"A repeated execute, observe, compare, and correct cycle that withholds completion until an attempted result satisfies explicit evidence or acceptance criteria.","definition_html":"<h2>Definition<\/h2>\n<p>A repeated execute, observe, compare, and correct cycle that withholds completion until an attempted result satisfies explicit evidence or <a href=\"\/glossary\/acceptance-criteria\" class=\"glossary-link\" title=\"Explicit conditions an outcome must satisfy before it can be accepted, promoted, or declared complete.\" data-glossary-slug=\"acceptance-criteria\">acceptance criteria<\/a>. A useful implementation separates the producer from the observation or grading step, records false alarms as well as caught defects, and limits retries when the same failure recurs.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A grader emits a judgment, an oracle establishes truth for a defined property, and a <a href=\"\/glossary\/verification-gate\" class=\"glossary-link\" title=\"A control point that blocks promotion until required evidence has been produced and validated.\" data-glossary-slug=\"verification-gate\">verification gate<\/a> enforces the admission decision. The verification loop connects those checks to another bounded attempt. Repeated self-critique without new evidence or an independent observation path is weaker than a verification loop.<\/p>\n<h2>Check your understanding<\/h2>\n<p>If the verifier rejects correct work too often, or the correction step can break correct work, another pass can reduce reliability instead of improving it.<\/p>\n","category":"evaluation-and-reliability","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["goal loop","grader loop","rubric loop"],"link_forms":[],"created_at":"2026-08-05T00:00:00-04:00","updated_at":"2026-08-05T00:00:00-04:00","related_terms":[{"slug":"grader","url":"https:\/\/darkfactory.dev\/glossary\/grader"},{"slug":"oracle","url":"https:\/\/darkfactory.dev\/glossary\/oracle"},{"slug":"verification-gate","url":"https:\/\/darkfactory.dev\/glossary\/verification-gate"},{"slug":"independent-verification","url":"https:\/\/darkfactory.dev\/glossary\/independent-verification"},{"slug":"loop-engineering","url":"https:\/\/darkfactory.dev\/glossary\/loop-engineering"}],"related_factory_areas":[{"slug":"verification","url":"https:\/\/darkfactory.dev\/factory\/verification"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Where Does Agent Reliability Come From?","url":"https:\/\/arxiv.org\/abs\/2607.17044"},{"title":"The Art of Loop Engineering: How to Build Agents That Improve Over Time","url":"https:\/\/www.youtube.com\/watch?v=jPPiZ22DY3g"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/vibe-coding","slug":"vibe-coding","term":"Vibe coding","definition":"Building software by prompting and accepting generated behavior with limited understanding or inspection of the underlying code.","definition_html":"<h2>Definition<\/h2>\n<p>Building software by prompting and accepting generated behavior with limited understanding or inspection of the underlying code.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Vibe coding minimizes code comprehension; <a href=\"\/glossary\/agentic-software-engineering\" class=\"glossary-link\" title=\"The discipline of designing software work so goal-directed AI agents can perform substantial engineering while humans retain product judgment, architecture, governance, and accountability.\" data-glossary-slug=\"agentic-software-engineering\">agentic engineering<\/a> uses agents while retaining deliberate specifications, verification, and ownership.<\/p>\n<h2>Check your understanding<\/h2>\n<p>The term is cultural and contested, so describe review, understanding, and accountability rather than using it as a quality judgment alone.<\/p>\n","category":"software-factory","definition_status":"contested","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"human-roles-expertise","url":"https:\/\/darkfactory.dev\/factory\/human-roles-expertise"}],"evidence":[{"title":"StrongDM: Software Factories and the Agentic Moment","url":"https:\/\/factory.strongdm.ai\/"},{"title":"AI-Generated Smells: Architecture Decay","url":"https:\/\/arxiv.org\/abs\/2605.02741"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/workflow","slug":"workflow","term":"Workflow","definition":"A defined sequence or graph of activities, decisions, states, and transitions used to achieve an outcome.","definition_html":"<h2>Definition<\/h2>\n<p>A defined sequence or graph of activities, decisions, states, and transitions used to achieve an outcome.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>A workflow prescribes process structure; an agent may choose actions dynamically within it.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Deterministic workflows and agentic decisions can coexist at different layers.<\/p>\n","category":"agents-and-automation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[{"slug":"control-graph","url":"https:\/\/darkfactory.dev\/glossary\/control-graph"},{"slug":"execution-graph","url":"https:\/\/darkfactory.dev\/glossary\/execution-graph"},{"slug":"directed-acyclic-graph","url":"https:\/\/darkfactory.dev\/glossary\/directed-acyclic-graph"},{"slug":"state-machine","url":"https:\/\/darkfactory.dev\/glossary\/state-machine"}],"related_factory_areas":[{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Codex Orchestration (Symphony)","url":"https:\/\/openai.com\/index\/open-source-codex-orchestration-symphony\/"},{"title":"A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents","url":"https:\/\/arxiv.org\/abs\/2605.20173"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/working-memory","slug":"working-memory","term":"Working memory","definition":"Short-lived task state actively used during a run, such as the current plan, recent observations, pending actions, and temporary summaries.","definition_html":"<h2>Definition<\/h2>\n<p>Short-lived task state actively used during a run, such as the current plan, recent observations, pending actions, and temporary summaries.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p>Working memory supports the current episode; <a href=\"\/glossary\/durable-memory\" class=\"glossary-link\" title=\"State intentionally retained across runs, such as verified facts, decisions, preferences, learned procedures, or persistent identity.\" data-glossary-slug=\"durable-memory\">durable memory<\/a> persists across sessions or tasks.<\/p>\n<h2>Check your understanding<\/h2>\n<p>A conversation transcript is one possible store but may be too noisy to function as useful working memory.<\/p>\n","category":"context-and-knowledge","definition_status":"working","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":["short-term memory"],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[{"slug":"context-memory-skills","url":"https:\/\/darkfactory.dev\/factory\/context-memory-skills"},{"slug":"orchestration-state","url":"https:\/\/darkfactory.dev\/factory\/orchestration-state"}],"evidence":[{"title":"Long-Running Agents","url":"https:\/\/addyosmani.com\/blog\/long-running-agents\/"},{"title":"Shepherd: A Runtime Substrate Empowering Meta-Agents with a Formalized Execution Trace","url":"https:\/\/arxiv.org\/abs\/2605.10913"}]},{"id":"https:\/\/darkfactory.dev\/glossary\/zero-shot-learning","slug":"zero-shot-learning","term":"Zero-shot learning","definition":"Performing a task or recognizing a category without task-specific labeled examples supplied for that use.","definition_html":"<h2>Definition<\/h2>\n<p>Zero-shot learning or prompting performs a task without task-specific examples supplied in the prompt or adaptation set.<\/p>\n<h2>Distinguish it from nearby terms<\/h2>\n<p><a href=\"\/glossary\/few-shot-prompting\" class=\"glossary-link\" title=\"Supplying a small set of worked examples in context to steer task behavior without updating model weights.\" data-glossary-slug=\"few-shot-prompting\">Few-shot prompting<\/a> supplies examples in context; zero-shot relies on instructions and prior learned capability.<\/p>\n<h2>Check your understanding<\/h2>\n<p>Ask whether any demonstrations, fine-tuning examples, or task-specific labels were actually provided.<\/p>\n","category":"inference-and-generation","definition_status":"stable","search_index":false,"search_index_reason":null,"search_reviewed_at":null,"aliases":[],"link_forms":[],"created_at":"2026-08-03T00:00:00-04:00","updated_at":"2026-08-03T00:00:00-04:00","related_terms":[],"related_factory_areas":[],"evidence":[{"title":"Google Machine Learning Glossary","url":"https:\/\/developers.google.com\/machine-learning\/glossary\/"}]}]}