63 terms
Definitions in this topic
- ProductionAI infrastructureAI infrastructure explained for buyers: what the stack contains, who provides it, and what you actually need to own.
- LLM foundationsAI21 LabsAI21 Labs explained as a language model developer and enterprise AI platform provider.
- LLM foundationsAlibaba Cloud AIAlibaba Cloud AI explained as a cloud platform for foundation models, model development, and hosted inference.
- ProductionAmazon BedrockAmazon Bedrock explained as AWS's managed platform for accessing and operating multiple foundation models.
- LLM foundationsAnthropicAnthropic explained as the company behind Claude, its developer platform, and the controls applications still need.
- LLM foundationsAudio language modelAudio language models explained: models that interpret or generate speech and sound, their representations, and how to evaluate them.
- LLM foundationsAutomatic speech recognitionAutomatic speech recognition explained: ASR, speech-to-text, streaming transcripts, accuracy limits, and its role in voice systems.
- ProductionAzure AI FoundryAzure AI Foundry, now Microsoft Foundry, explained as Microsoft's platform for building and operating AI applications.
- ProductionBring Your Own ModelBring your own model defined, including deployment patterns, benefits, and operating responsibilities.
- ProductionCerebrasCerebras explained as an AI compute and inference provider using wafer-scale processor architecture.
- LLM foundationsCohereCohere explained as an enterprise AI provider for generation, retrieval, reranking, and multilingual systems.
- LLM foundationsDeepSeekDeepSeek explained as an AI model developer, including open models, hosted access, and production evaluation criteria.
- LLM foundationsFine-tuningFine-tuning explained: what it actually changes, and why most teams should try RAG or better context first.
- ProductionFireworks AIFireworks AI explained as managed infrastructure for serving, customizing, and scaling generative models.
- LLM foundationsFrontier modelFrontier model explained: what puts a model at the frontier, versus foundation models, and why the label keeps moving.
- LLM foundationsGenerative AIGenerative AI explained for business readers: what it is, how it differs from agentic AI, and where the value shows up.
- LLM foundationsGoogle DeepMindGoogle DeepMind explained: its research role, relationship with Google, and how businesses access its models.
- ProductionGoogle Vertex AIGoogle Vertex AI explained as Google Cloud's managed platform for models, agents, evaluation, and AI operations.
- ProductionGroqGroq explained as an AI inference provider built around LPU hardware, and how it differs from xAI's Grok.
- ProductionHosted InferenceHosted inference defined, including its operational benefits, tradeoffs, and production evaluation criteria.
- ProductionHugging FaceHugging Face explained as the AI platform for models, datasets, open-source libraries, and hosted inference.
- LLM foundationsInferenceInference explained: what actually happens per API call, why it costs money every time, and how to keep it fast.
- ProductionInference ProviderInference providers run AI models behind an API. How they differ from LLM providers, how to compare them, and what we learned running one.
- LLM foundationsLLMWhat an LLM actually is from a buyer's seat: what it can carry in production, and where it needs a system built around it.
- ProductionLLM ProviderLLM provider meaning: model labs, cloud platforms and inference hosts, plus which providers Soba Labs runs in production and why we use more than one.
- LLM foundationsMeta AIMeta AI explained as a model developer, including open weights, hosting choices, and deployment responsibility.
- LLM foundationsMiniMaxMiniMax explained as a multimodal AI developer providing models and services across text, audio, image, and video.
- LLM foundationsMistral AIMistral AI explained as a European model provider offering hosted APIs and selected open weight models.
- LLM foundationsMixture of ExpertsMixture of Experts explained: how sparse routing gives language models more capacity without running every parameter.
- ProductionModel distillationModel distillation explained: how a small student learns from a big teacher, and when it beats fine-tuning or quantization.
- ProductionModel routingModel routing explained: choosing an LLM per request to balance capability, latency, cost, privacy, and availability.
- LLM foundationsMoonshot AIMoonshot AI explained as the developer of Kimi models, with practical criteria for production evaluation.
- ProductionMulti-Provider AIMulti-provider AI defined, including routing, resilience, portability, and the operational cost of multiple vendors.
- LLM foundationsMultimodal LLMMultimodal LLMs explained: models that combine language reasoning with images, audio, video, or other data types.
- ProductionNVIDIA AINVIDIA AI explained across GPUs, inference software, models, and enterprise deployment infrastructure.
- LLM foundationsOpen-weight modelOpen-weight models explained: downloadable parameters, licenses, and the open-source distinction.
- LLM foundationsOpenAIOpenAI explained as an AI model provider, including its APIs, platform role, and production considerations.
- ProductionOpenRouterOpenRouter explained as a shared API for accessing, comparing, and routing requests across many AI models.
- LLM foundationsPost-trainingPost-training explained: how fine-tuning, preference optimization and RL turn a pre-trained model into a useful one.
- ProductionQuantizationQuantization explained: what fewer bits per weight buy you, what quality they cost, and when a compressed model is the right call.
- LLM foundationsReasoning effortReasoning effort explained: the levels OpenAI and Anthropic accept, what each one costs, and the effort settings we have seen fail silently.
- LLM foundationsReasoning modelsReasoning models explained: what extended thinking actually buys you, and when the extra latency is worth paying.
- ProductionReplicateReplicate explained as a hosted API platform for running community and publisher-provided AI models.
- ProductionServerless InferenceServerless inference defined, including autoscaling benefits and the latency, capacity, and cost tradeoffs.
- LLM foundationsSmall Language ModelSmall language model explained: when a compact model is faster, cheaper, and more private than a general LLM.
- LLM foundationsSpeaker diarizationSpeaker diarization explained: identifying who spoke when, how it differs from speaker identification, and common accuracy limits.
- LLM foundationsSpeaker identificationSpeaker identification explained: matching voices to enrolled identities, how it differs from diarization, and why confidence needs controls.
- LLM foundationsSpeculative DecodingSpeculative decoding explained: how draft and target models verify several tokens at once to reduce LLM latency.
- LLM foundationsSpeech-to-speech modelSpeech-to-speech models explained: direct audio input and output, how they differ from cascaded voice systems, and where controls remain essential.
- ProductionStreaming inferenceStreaming inference explained: incremental model output, lower perceived latency, partial-result handling, cancellation, and reliability.
- LLM foundationsTemperatureTemperature explained: how the sampling parameter trades determinism for variety, and how to set it per task.
- LLM foundationsTest-Time ComputeTest-time compute explained: spending more inference budget on reasoning, sampling, verification, and answer selection.
- LLM foundationsText-to-speechText-to-speech explained: how TTS generates audio, what affects quality, and which consent and production controls teams need.
- ProductionTogether AITogether AI explained as an infrastructure provider for hosted inference, training, and deployment of open models.
- LLM foundationsTokenToken explained: what a token actually is, why it is not a word, and where the count quietly adds up.
- LLM foundationsTokenizationTokenization explained: how language models convert text into tokens and why token counts affect context, cost, and output.
- LLM foundationsTransformer architectureTransformer architecture explained: how attention connects tokens and powers modern language models.
- LLM foundationsVision-language modelVision-language models explained: combining visual inputs with language instructions for analysis, extraction, and agents.
- ProductionVoice activity detectionVoice activity detection explained: how VAD finds speech in audio, supports turn boundaries, and differs from semantic completion.
- LLM foundationsVoice cloningVoice cloning explained: how models reproduce a person's voice, legitimate uses, consent requirements, and impersonation risks.
- LLM foundationsWorld modelWorld models explained: how they differ from LLMs, why labs are building them, and what they unlock for agents and robotics.
- LLM foundationsxAIxAI explained as the company behind Grok, including API access and what businesses should evaluate.
- LLM foundationsZ.aiZ.ai explained as a provider of GLM foundation models and developer services for AI applications.