Tech, translated.
659 terms in plain English
Look up any term on a resume or JD. See what it means, what it says about seniority, and exactly what to ask.
Just added
The 121 that tell you the most about a candidate.
AI research roles. Research Scientists invent new methods and usually publish papers; Applied Scientists (a common Amazon and Microsoft title) adapt research to real products; Research Engineers build the code and infrastructure that research runs on. Most research scientist roles expect a PhD or a strong publication record.
An open-source framework for spreading Python and AI workloads — data processing, model training, tuning and serving — across many machines. Widely used by companies that train or run their own models at scale. Anyscale is the company behind it.
Google's framework for high-performance machine learning and scientific computing, used heavily at Google DeepMind and for training large models on Google's TPUs. Flax is the most common library for building neural networks with it. An alternative to PyTorch, which is more widely used elsewhere.
Writing the low-level code that runs directly on GPUs — the small programs (kernels) that do the core maths of AI models. CUDA is NVIDIA's platform, Triton is a newer, easier language for the same job, and ROCm is AMD's equivalent. Hand-tuned kernels can make training or inference several times faster.
Making AI models answer faster and more cheaply on the same hardware. Techniques include serving many requests together (continuous batching), using a small model to draft words a big model then checks (speculative decoding), managing memory more efficiently, and compressing the model. vLLM, SGLang, TensorRT-LLM and llama.cpp are common tools.
Research into what is actually happening inside a neural network: finding which internal parts represent which concepts and how they combine to produce an answer. Done mainly at AI labs and safety organisations to understand, debug and make models safer. Different from explainable AI, which explains a model's individual decisions from the outside.
Training a model with reinforcement learning on tasks where the answer can be checked automatically — did the code pass its tests, is the maths answer correct — instead of relying on human ratings. This is how reasoning and coding models are mostly trained today. Building 'RL environments' (realistic practice tasks with automatic grading) has become its own fast-growing job.
Everything done to a model after its initial large-scale training to make it useful and safe: teaching it to follow instructions with curated examples (supervised fine-tuning, or SFT), then refining its behaviour with feedback methods such as RLHF, DPO and reinforcement learning. A core research team at every AI lab, and the stage that gives each model its personality and skills.
The technology built into browsers and phones for live audio, video and data between users. It powers video calls, telehealth consultations, live tutoring and voice AI agents. Making it reliable over poor networks and corporate firewalls needs extra relay and media servers; companies such as LiveKit and Agora sell that infrastructure.
Very fast analytical databases that answer questions over billions of rows in under a second, often on data that arrived moments ago. Used behind live dashboards, product analytics and observability tools. ClickHouse is the best known; Apache Druid, Apache Pinot and StarRocks do similar jobs.
A NoSQL database built to spread huge amounts of data across many servers and keep accepting writes even when some servers, or a whole data centre, fail. Used for messaging, activity feeds, IoT and other workloads with enormous write volumes. Data has to be designed around the exact questions the application will ask, which is very different from a traditional database. ScyllaDB is a faster compatible alternative.
A way of storing shared data so several people or devices can edit it at the same time, even offline, and every copy automatically ends up identical without a central server settling conflicts. One of the main techniques behind live multi-user editing in collaborative apps. Yjs and Automerge are popular libraries, and 'local-first' apps are built on the idea.
Protecting data even while it's being processed, by running code inside a sealed-off area of the chip (a trusted execution environment, or enclave) that even the cloud provider or server administrator can't see into. Used for highly sensitive data, analytics shared between companies and, increasingly, running AI models on private data.
A Linux technology that lets engineers run small, safe programs inside the operating system itself to watch or control network traffic, security events and performance, without changing application code. Cilium (Kubernetes networking) and many modern monitoring and security tools are built on it.
New encryption methods designed to stay secure even against a future large quantum computer, which could break much of today's encryption. The US standards body NIST published the first standards in 2024, and organisations are now working through the long job of finding and replacing old encryption everywhere it is used.
Computers that use qubits, which follow the rules of quantum physics, to tackle certain problems (simulating molecules, some optimisation and cryptography problems) that ordinary computers can't solve in a reasonable time. Today's machines are still small and error-prone, so most work is research, algorithm design and error correction. IBM's Qiskit is the most common programming toolkit.
AI that controls machines in the physical world, such as robot arms, humanoid robots, drones and self-driving vehicles, using large models that turn camera input and instructions directly into movement. World models, which simulate how the physical world behaves, are used to train and test these systems.
The practices for running LLM-based products reliably once they're live: versioning prompts and models, logging every request, monitoring cost, speed and answer quality, and running evaluations before each change ships. The LLM-era cousin of MLOps.
Training a model across many devices or organisations without moving their data to one place. Each device learns from its own data and sends back only model updates. Used for phone keyboards, and for hospitals or banks that cannot legally share raw records.
Combining several teams' GraphQL services into one unified API (a 'supergraph'), so frontend apps query a single endpoint while each team owns its own part. Apollo is the main vendor.
Elixir is a language built on Erlang's platform (BEAM), designed for systems handling huge numbers of simultaneous connections that must not go down, such as chat, telecoms and real-time apps. Phoenix is its web framework.
Tools for running long, multi-step business processes (orders, onboarding, payments) that must survive crashes and resume where they left off. Temporal is the best-known; AWS Step Functions is Amazon's managed version.
One running system serving many customer organisations at once while keeping each one's data strictly separate — the basis of most B2B SaaS products.
An approach to designing complex software around the business itself: shared language with business experts, and clear boundaries ('bounded contexts') between areas like billing and shipping. Often used to decide how to split microservices.
A Webpack feature (now in other bundlers too) that lets separately built and deployed frontend apps share code and load each other's parts at runtime. It's the most common way micro-frontends are wired together.
Technologies for drawing fast 2D and 3D graphics in the browser using the graphics card. Three.js is the most popular library on top of WebGL; WebGPU is its newer successor.
A way to run heavy JavaScript in the background of a web page, so expensive work (parsing big files, image processing) doesn't freeze the screen.
The first, enormously expensive stage of building a large model: training it on a huge slice of the internet, books and code so it learns language and general knowledge. Fine-tuning and RLHF come afterwards to make it helpful and safe. The raw result is called a base model.
The large IBM machines and the 1960s-era COBOL language that still process most bank card payments, insurance claims and government benefits. The systems are old but critical, and there are fewer and fewer people who know them.
A written, usually machine-checked agreement between the team that produces data and the teams that use it, covering which fields exist, what they mean and how fresh they are, so an upstream change can't silently break reports downstream.
Whether a model's mistakes fall unevenly on particular groups of people, usually because the training data carried a historic pattern. Directly regulated where models touch hiring, lending or insurance.
Splitting one model's training across many machines because it will not fit, or will not finish, on a single one. Splitting the data is the common case; splitting the model itself is far harder.
The specialised chips model training and inference run on. They perform many small calculations at once, which is what makes training practical, and they are the dominant cost line in most AI work.
Designing a system so that when one part fails, the rest keeps working with reduced function instead of the whole thing going down.
The internal reliability target a team holds itself to, and the specific measurement it is judged by. Distinct from an SLA, which is the contractual promise made to a customer.
Code that fetches a list, then quietly runs one more database query for every item in it — fine with ten rows, crippling with ten thousand.
A system where a change reaches every copy of the data soon, but not instantly — so for a moment, two users can see different answers.
A shared store of the prepared inputs models use, so the same calculation feeds both model training and the live system.
Two different ways of being right. Precision asks how many of the things the model flagged were real. Recall asks how many of the real things it caught. Improving one usually costs the other.
Designing an app to work fully without a connection and reconcile changes when it returns — rather than treating offline as an error state.
Sharing business logic across iOS and Android in Kotlin while keeping each platform's native UI — a middle path between fully native and fully cross-platform.
Rendering a page as static HTML and hydrating only the interactive regions, so a mostly-static page doesn't pay for a full JavaScript app.
Two processes each holding something the other needs, so neither can continue — the system stops without crashing.
A bug where the outcome depends on the timing of two things happening at once — the kind that passes every test and fails in production under load.
A vendor-neutral standard for emitting traces, metrics and logs, so you can change observability tooling without re-instrumenting every service.
Following a single request as it passes through many services, so you can see where the time went and which hop failed.
Deliberately slowing or rejecting incoming work when a system is overloaded, so it degrades predictably instead of collapsing.
A way to keep a multi-service operation consistent without a single database transaction — each step has a defined undo, run when a later step fails.
Storing every change as an immutable event and deriving current state by replaying them, instead of overwriting rows in place.
Splitting the code paths that change data from the ones that read it, so each can be designed and scaled for its own very different job.
The unit economics of running a model — cost scales with text in and out, so a feature that looks cheap in testing can be ruinous at scale.
How quickly a model starts responding, as distinct from how long the full answer takes. Streaming the first words fast makes an application feel responsive even when total time is unchanged.
Training a small, cheap model to imitate a large expensive one, so you get most of the quality at a fraction of the cost and latency.
Techniques for adapting a large model by training a small number of extra parameters instead of retraining the whole thing — far cheaper, and the usual way fine-tuning is actually done.
The infrastructure that runs a model and answers requests — handling batching, GPU memory and concurrency so that responses stay fast under load.
A second pass that reorders search results by relevance using a slower, more accurate model — you retrieve fifty candidates cheaply, then carefully rank the top few.
The practice of managing cloud spend as an engineering concern — attributing costs to teams and making the people who create spend able to see it.
A branching approach where everyone merges small changes into the main branch at least daily, instead of maintaining long-running feature branches.
Four research-backed measures of delivery performance: deployment frequency, lead time for changes, change failure rate, and time to restore service.
A self-service layer that lets product engineers provision environments and ship services without filing tickets or learning the whole infrastructure stack.
Treating a company's own engineers as users — measuring and improving how fast and how painlessly they can build, test and ship.
The system of certificate authorities, keys and trust chains that makes it possible to verify that a server or user is who it claims to be.
A setup where both sides of a connection prove who they are with certificates, not just the server — common between internal services in zero-trust architectures.
Protecting everything that goes into building and shipping software — dependencies, build systems, signing keys — rather than just the finished application.
A machine-readable inventory of every component and dependency inside a piece of software — increasingly required by regulators and enterprise buyers.
An approach where each business domain owns and publishes its own data as a product, instead of one central data team owning everything.
A technique that streams every change made in a database to other systems as it happens, instead of re-copying whole tables on a schedule.
A stream-processing engine for computing over data as it arrives, continuously, rather than in scheduled batches.
Sits between Data Engineer and Data Analyst — takes raw ingested data and models it into clean, documented, tested tables that analysts can trust. Usually dbt-centric.
Testing that a product is actually usable by people with disabilities — screen readers, keyboard-only navigation, color contrast — using both automated scanners and real manual/assistive-technology checks, since automated tools alone catch only a fraction of real issues.
Moving testing earlier in the development process — writing tests alongside (or before) the code, rather than only after a feature is complete — to catch bugs when they're cheapest to fix.
Roughly calculating a system's expected scale — storage needed, requests per second, bandwidth — early in a design, to sanity-check whether an approach can actually work before building it.
Algorithms (like Paxos and Raft) that let multiple machines in a distributed system agree on a single value or decision, even if some machines fail or messages get delayed — the foundation under things like leader election.
A technique for distributing data across multiple servers so that when a server is added or removed, only a small fraction of data needs to move — instead of nearly everything reshuffling.
A system's ability to keep working correctly even when part of it fails — a server crashes, a network link drops — instead of the whole system going down.
Systems made up of multiple independent computers that coordinate over a network to work as one, instead of running on a single machine — introducing challenges like partial failures and network delays that don't exist on one box.
A guideline for how a healthy test suite should be shaped: many fast, cheap unit tests at the base, fewer integration tests in the middle, and very few slow, expensive end-to-end tests at the top.
Testing that verifies two services (like a frontend and an API, or two microservices) agree on the shape of the data they exchange, catching breaking changes before they reach production — without needing to spin up both full systems together.
A popular tool for automating mobile app build, testing, and release tasks (App Store/Play Store submission, code signing, screenshots) that would otherwise be manual and error-prone.
Systematic tests that measure how well an AI system performs on a task, used to catch regressions and compare model or prompt changes objectively instead of eyeballing a few examples.
Shrinking a model's size and memory footprint by reducing the precision of its internal numbers, trading a small amount of accuracy for speed and lower cost.
An attack where malicious text (hidden in a document, webpage, or user input) tricks an AI system into ignoring its original instructions and doing something unintended.
A training technique where human reviewers rank AI outputs, and the model is further trained to produce more of what humans preferred — a key step in making raw models helpful and safe.
When a deployed model's real-world performance degrades over time because the data it now sees no longer matches the data it was trained on — the world changed, but the model didn't.
A machine learning approach where a model learns by taking actions in an environment and getting rewards or penalties for the outcomes, gradually improving its strategy through trial and error — rather than learning from a fixed labeled dataset.
Practices and tooling for reliably deploying, monitoring, and maintaining machine learning models in production (the ML equivalent of DevOps).
A browser security feature that lets a site declare which sources of scripts, styles, and other content are allowed to load, blocking a large class of cross-site scripting (XSS) attacks.
The practice of proactively identifying how a system could be attacked and what its weak points are, before it ships — rather than discovering vulnerabilities only after an incident.
Securely storing and controlling access to sensitive credentials (API keys, passwords, certificates) instead of hardcoding them into source code or config files.
A security model that assumes no user or system should be automatically trusted, even inside the network — every request is verified.
SRE terminology for manual, repetitive operational work that doesn't add lasting value and scales linearly with system growth — a target for automation, not a badge of hard work.
The practice of deliberately injecting failures into a production or production-like system (killing a server, cutting network access) to test how well it holds up, before a real failure happens unplanned.
Standard incident metrics: MTTA is how long it takes someone to acknowledge an alert; MTTR is how long it takes to actually resolve the underlying issue once it's known.
A written review conducted after an incident to understand what went wrong and prevent it from recurring, ideally 'blameless' — focused on the system and process, not blaming an individual.
The structured process a team follows when something breaks in production — detecting, communicating, mitigating, and resolving the issue.
The amount of allowed unreliability (e.g. the 0.1% downtime under a 99.9% SLA) a team is permitted to 'spend' before it must pause new feature work and focus on stability.
A team's planned process for restoring systems and data after a major failure (a data center outage, data corruption, a botched deployment) — including how much data loss and downtime is acceptable.
Using more than one cloud provider (e.g. AWS and GCP together) rather than committing entirely to one, often to avoid vendor lock-in or use each provider's specific strengths.
Deployment strategies for releasing new code with less risk: blue-green runs two full identical environments and switches traffic all at once; canary rolls the new version out to a small slice of traffic first before going wider.
Infrastructure that manages communication between microservices — routing, retries, encryption, and observability — without each individual service having to implement that logic itself.
An approach where a Git repository is the single source of truth for what infrastructure/deployments should look like — a tool automatically syncs the live system to match whatever's committed to Git.
The problem of knowing when cached data has gone stale and needs to be refreshed or removed — famously one of the hardest problems in computer science because it's easy to get subtly wrong.
A database optimized for storing and querying data points tagged with timestamps at high volume — metrics, sensor readings, logs — with built-in support for time-based aggregation and retention.
A database that stores data as nodes and the relationships between them, optimized for queries about how things connect — like social networks, recommendation engines, or fraud detection — rather than rows and columns.
Reusing a shared set of open database connections across requests instead of opening and closing a new one every time, which is slow and resource-heavy at scale.
A principle stating a distributed database can't simultaneously guarantee Consistency (everyone sees the same data), Availability (every request gets a response), and Partition tolerance (it keeps working despite network failures) — it has to trade off between them.
Splitting a database's data across multiple separate database servers (shards), so no single server has to hold or serve all the data.
The process of rewriting a database query or adjusting a database's structure so a slow query runs faster, often by reducing how much data it has to scan.
Techniques for a program to handle multiple tasks that overlap in time — either genuinely in parallel (multithreading) or by not blocking while waiting on slow operations like network calls (async/await).
An architecture where services communicate by publishing and reacting to events (things that happened) rather than calling each other directly, letting services stay decoupled from one another.
A property of an operation where running it multiple times has the same effect as running it once — important for safely retrying requests (like a payment) without accidentally duplicating the result.
A pattern that stops a service from repeatedly calling another service that's already failing, temporarily 'breaking the circuit' to prevent one failure from cascading into a larger outage.
A high-performance framework (from Google) for services to call each other directly, often faster than REST for service-to-service communication, commonly used inside microservices architectures.
A low-level format that lets code written in languages like C++, Rust, or Go run in the browser at near-native speed, used for performance-critical tasks JavaScript handles poorly (video editing, games, simulations).
A script that runs in the background of a browser tab, separate from the page itself, enabling offline support, caching, and push notifications for web apps.
An architecture where a large web app is split into smaller, independently built and deployed frontend pieces — owned by different teams — instead of one single frontend codebase.
A documented, shared set of design rules, components, and guidelines (colors, spacing, typography, components) that keeps a product visually and functionally consistent across teams.
A set of metrics Google uses to measure real-world user experience on a page — how fast it loads (LCP), how stable the layout is as it loads (CLS), and how responsive it feels to interact with (INP).
A newer React feature that lets some components render entirely on the server and send only the result to the browser, reducing the amount of JavaScript the client has to download and run.
The process of a server-rendered page 'waking up' in the browser — JavaScript attaches to the already-rendered HTML and makes it interactive, instead of building the page from scratch.
A set of HTML attributes that give assistive technology (like screen readers) extra information about custom interactive elements that plain HTML alone doesn't convey.