What is a GPU, actually?
Before writing any code, it helps to have a mental model of the hardware we're programming. A GPU is fundamentally different from a CPU:
- A CPU has a few very powerful cores (4–32), each optimised for executing complex sequential logic quickly, with large out-of-order execution windows and branch predictors.
- A GPU has thousands of much simpler cores (a modern GPU has 5,000–20,000 "shader processors"), each optimised for doing simple math on independent data in parallel.
This makes GPUs ideal for graphics: shading a triangle means running the same program on potentially millions of pixels simultaneously, with no dependencies between them. A CPU would process them one-by-one; a GPU does them all at once.
The GPU also has its own dedicated memory — VRAM (Video RAM). It's separate from your system RAM. When you upload geometry or textures, you're copying data from system RAM across the PCIe bus into VRAM. The GPU can read VRAM very quickly, but transferring between CPU and GPU memory is expensive — so you want to minimise those transfers in your render loop.
The rendering pipeline
To draw anything on screen, your data passes through a fixed sequence of stages called the rendering pipeline. Some stages are programmable — you write code that runs on the GPU (called shaders). Others are fixed-function — the hardware does them automatically, you just configure them.
Here's the pipeline for drawing a triangle:
Let's go through the key stages:
Coordinate systems and NDC
This trips up almost everyone. OpenGL uses several coordinate systems, and understanding how they chain together is fundamental to 3D rendering.
The one you need to know to render a triangle is NDC — Normalized Device Coordinates. After the vertex shader runs, the GPU expects vertex positions to be in NDC space: a cube from -1 to +1 on all axes. Anything outside that cube is off-screen and clipped.
For our first triangle, we'll place vertices directly in NDC so the vertex shader can be a simple passthrough. In a real 3D program, you'd transform from "world space" (where your scene objects live) through "view space" (relative to the camera) and then through a projection matrix into clip space. But NDC first, math later.
Setting up: window + OpenGL context
Before calling any OpenGL function, you need two things: a window to render into, and an OpenGL context — the object that represents your connection to the GPU driver.
Creating a window is platform-specific (Win32 on Windows, X11/Wayland on Linux, Cocoa on macOS). Doing it manually for each platform is painful. GLFW is a small library that handles all of this portably.
Then there's a second problem: OpenGL function pointers aren't linked at compile time — they're loaded at runtime from the driver. Doing this manually for the 300+ functions in OpenGL 3.3 would be absurd. GLAD is a loader generator that handles this.
/* Dependencies you need:
- GLFW: window + context creation (https://www.glfw.org)
- GLAD: OpenGL function loader (https://glad.dav1d.de)
Generate for: OpenGL 3.3, Core profile, C/C++ language
Compile: gcc main.c glad.c -lglfw -lGL -lm -o triangle
(On macOS: -framework OpenGL instead of -lGL)
*/
#include "glad/glad.h" // must be included BEFORE glfw
#include
#include
#include
/* Called automatically when the window is resized */
void framebuffer_size_callback(GLFWwindow *window, int width, int height) {
/* Tell OpenGL the new rendering region size */
glViewport(0, 0, width, height);
}
int main(void) {
/* ── 1. Initialise GLFW ── */
if (!glfwInit()) {
fprintf(stderr, "Failed to initialise GLFW\n");
return -1;
}
/* Tell GLFW which OpenGL version to use and request Core Profile.
Core Profile removes all the deprecated "legacy" OpenGL stuff —
forces you to do things the modern way. */
glfwWindowHint(GLFW_CONTEXT_VERSION_MAJOR, 3);
glfwWindowHint(GLFW_CONTEXT_VERSION_MINOR, 3);
glfwWindowHint(GLFW_OPENGL_PROFILE, GLFW_OPENGL_CORE_PROFILE);
/* macOS needs this extra hint — Apple's OpenGL support is frozen at 4.1
but uses a forward-compat context to expose it. */
#ifdef __APPLE__
glfwWindowHint(GLFW_OPENGL_FORWARD_COMPAT, GL_TRUE);
#endif
/* ── 2. Create window + context ── */
GLFWwindow *window = glfwCreateWindow(800, 600, "Hello Triangle", NULL, NULL);
if (!window) {
fprintf(stderr, "Failed to create GLFW window\n");
glfwTerminate();
return -1;
}
/* Make this window's GL context the current one on this thread.
All subsequent OpenGL calls operate on this context. */
glfwMakeContextCurrent(window);
glfwSetFramebufferSizeCallback(window, framebuffer_size_callback);
/* ── 3. Load OpenGL function pointers via GLAD ── */
/* glfwGetProcAddress is GLFW's platform-specific function pointer
loader — GLAD calls it for every function it needs to load. */
if (!gladLoadGLLoader((GLADloadproc)glfwGetProcAddress)) {
fprintf(stderr, "Failed to initialise GLAD\n");
return -1;
}
/* From this point, all gl*() functions are available.
We can check what we got: */
printf("OpenGL version: %s\n", glGetString(GL_VERSION));
printf("GPU: %s\n", glGetString(GL_RENDERER));
// ... (shader setup, VBO/VAO, render loop) ...
glfwDestroyWindow(window);
glfwTerminate();
return 0;
}
Getting geometry onto the GPU (VBO + VAO)
Your triangle's vertices start as a float array in system RAM. The GPU can't see your process's memory — you have to explicitly upload the data to VRAM and tell the GPU how to interpret it. OpenGL uses two objects for this:
- VBO (Vertex Buffer Object) — a chunk of GPU memory that stores raw vertex data. Think of it as
mallocon the GPU side. - VAO (Vertex Array Object) — remembers how your VBO data is structured (where each attribute starts, how many floats per vertex, etc). Think of it as a "format descriptor" or "schema" for a VBO.
The split exists for a good reason: you might have multiple VBOs with the same layout, or you might want to switch between different data layouts efficiently. The VAO captures all the state so you can bind it once and the GPU knows exactly how to read the data.
/* Our triangle: 3 vertices, each with an XYZ position (3 floats = 12 bytes) */
float vertices[] = {
/* x y z */
-0.5f, -0.5f, 0.0f, /* bottom-left */
0.5f, -0.5f, 0.0f, /* bottom-right */
0.0f, 0.5f, 0.0f, /* top-center */
};
GLuint vao, vbo;
/* ── 1. Create and bind the VAO first ──
Everything we configure below gets recorded into this VAO. */
glGenVertexArrays(1, &vao);
glBindVertexArray(vao);
/* ── 2. Create a VBO and upload data ── */
glGenBuffers(1, &vbo);
glBindBuffer(GL_ARRAY_BUFFER, vbo); /* bind to GL_ARRAY_BUFFER slot */
glBufferData(
GL_ARRAY_BUFFER, /* target: the slot we just bound to */
sizeof(vertices), /* size in bytes: 9 floats × 4 bytes = 36 */
vertices, /* the actual data to upload */
GL_STATIC_DRAW /* usage hint: data won't change after upload.
GL_DYNAMIC_DRAW = updated frequently.
GL_STREAM_DRAW = updated every frame.
These are hints to the driver about where
to store the buffer (fast VRAM vs slower). */
);
/* ── 3. Tell the GPU how to interpret the buffer data ──
This is the crucial step that often confuses beginners.
We're describing the memory layout of ONE vertex to the GPU. */
glVertexAttribPointer(
0, /* attribute index: this is "attribute 0".
In the vertex shader we'll write: layout(location=0) */
3, /* number of components: 3 floats per vertex (X, Y, Z) */
GL_FLOAT, /* data type of each component */
GL_FALSE, /* normalise? (only relevant for integer types) */
3 * sizeof(float), /* stride: how many bytes between the START of one
vertex and the start of the next. Here: 12 bytes.
If we had position + color: 6 floats = 24 bytes stride. */
(void*)0 /* offset: how many bytes from the start of the buffer
where this attribute's data begins. 0 = right at the start.
If color came after position: (void*)(3 * sizeof(float)) */
);
/* Enable attribute index 0 (disabled by default) */
glEnableVertexAttribArray(0);
/* Unbind the VAO — good practice to avoid accidentally modifying it later */
glBindVertexArray(0);
{ float x, y, z; float r, g, b; } — 24 bytes total. Position starts at byte 0, color starts at byte 12. You'd call glVertexAttribPointer twice: once for position (offset 0, stride 24) and once for color (offset 12, stride 24). The stride tells the GPU how to jump from one vertex to the next; the offset tells it where within each vertex to find this particular attribute.
Writing your first shaders (GLSL)
Shaders are programs written in GLSL (OpenGL Shading Language) — a C-like language that compiles and runs on the GPU. You provide them as source code strings at runtime; the driver compiles them when you call the appropriate OpenGL functions.
For our triangle we need exactly two shaders:
// vertex shader
/* vertex shader source — stored as a string in your C code */
#version 330 core
/* ↑ GLSL version 330 corresponds to OpenGL 3.3.
"core" means we're using Core Profile features only. */
/* Declare an input attribute at location 0.
This matches the '0' we passed to glVertexAttribPointer.
The GPU will feed each vertex's XYZ data into 'aPos'. */
layout (location = 0) in vec3 aPos;
void main() {
/* gl_Position is a built-in output — the GPU reads it after
the vertex shader to know where to place this vertex on screen.
It expects a vec4 (homogeneous coordinates: x, y, z, w).
For now, w=1.0 means "no perspective division" — our triangle
stays exactly where we placed it in NDC. */
gl_Position = vec4(aPos.x, aPos.y, aPos.z, 1.0);
/* Equivalently: gl_Position = vec4(aPos, 1.0);
vec3 can be swizzled into vec4 this way. */
}
This vertex shader does almost nothing — it passes the vertex position straight through. In a real program, you'd multiply by a model-view-projection matrix here to handle 3D transforms and camera perspective. But for NDC coordinates, passthrough is exactly right.
// fragment shader
/* fragment shader source */
#version 330 core
/* Declare an output variable — this is the color the shader writes.
'FragColor' is a name we choose; OpenGL routes location 0's output
to the default framebuffer (the screen). */
out vec4 FragColor;
void main() {
/* vec4(R, G, B, A) — all values 0.0 to 1.0.
This gives every pixel of our triangle a solid orange color. */
FragColor = vec4(1.0f, 0.5f, 0.1f, 1.0f);
/* Try changing these values to see different colors:
vec4(1.0, 0.0, 0.0, 1.0) = red
vec4(0.0, 1.0, 0.0, 1.0) = green
vec4(0.2, 0.6, 1.0, 1.0) = sky blue */
}
vec2, vec3, vec4 for floats; ivec2/3/4 for ints; bvec2/3/4 for bools. There are also matrix types: mat2, mat3, mat4. You can swizzle components with .x .y .z .w or .r .g .b .a (they're identical) and even chain them: v.xyz extracts a vec3, v.zyx reverses the order.
Compiling and linking shaders
Shaders need to be compiled at runtime by the GPU driver. The process mirrors traditional compilation: compile each shader individually, then link them into a shader program that the GPU executes.
/* Helper: compile a single shader and return its ID.
Prints the error log if compilation fails.
type: GL_VERTEX_SHADER or GL_FRAGMENT_SHADER */
GLuint compile_shader(GLenum type, const char *source) {
GLuint shader = glCreateShader(type);
glShaderSource(
shader,
1, /* number of source strings */
&source, /* array of source string pointers */
NULL /* array of string lengths (NULL = null-terminated) */
);
glCompileShader(shader);
/* Check for compilation errors */
GLint success;
glGetShaderiv(shader, GL_COMPILE_STATUS, &success);
if (!success) {
char log[1024];
glGetShaderInfoLog(shader, sizeof(log), NULL, log);
fprintf(stderr, "Shader compile error (%s):\n%s\n",
type == GL_VERTEX_SHADER ? "vertex" : "fragment", log);
glDeleteShader(shader);
return 0; /* 0 is not a valid shader ID */
}
return shader;
}
/* Helper: link a vertex and fragment shader into a complete program */
GLuint create_shader_program(const char *vert_src, const char *frag_src) {
GLuint vert = compile_shader(GL_VERTEX_SHADER, vert_src);
GLuint frag = compile_shader(GL_FRAGMENT_SHADER, frag_src);
if (!vert || !frag) return 0;
GLuint program = glCreateProgram();
glAttachShader(program, vert);
glAttachShader(program, frag);
glLinkProgram(program);
/* Check for link errors */
GLint success;
glGetProgramiv(program, GL_LINK_STATUS, &success);
if (!success) {
char log[1024];
glGetProgramInfoLog(program, sizeof(log), NULL, log);
fprintf(stderr, "Shader link error:\n%s\n", log);
glDeleteProgram(program);
return 0;
}
/* Individual shaders are no longer needed once linked into a program.
The program holds its own compiled copy. */
glDeleteShader(vert);
glDeleteShader(frag);
return program;
}
GL_COMPILE_STATUS and print the info log. The error messages are actually quite good — they'll tell you the exact line number and a description of the problem, much like a C compiler would.
The render loop
With the setup done, the actual rendering is a tight loop: clear the screen, bind the shader program, bind the geometry, draw it, swap the buffers, repeat until the window closes.
while (!glfwWindowShouldClose(window)) {
/* ── 1. Process window events (keyboard, mouse, resize, close) ──
Must be called every frame or the OS thinks the program is frozen. */
glfwPollEvents();
/* Handle Escape key to close */
if (glfwGetKey(window, GLFW_KEY_ESCAPE) == GLFW_PRESS)
glfwSetWindowShouldClose(window, 1);
/* ── 2. Clear the framebuffer ──
glClearColor sets the color to clear to (R, G, B, A).
glClear actually performs the clear on the specified buffers. */
glClearColor(0.08f, 0.08f, 0.10f, 1.0f); /* near-black background */
glClear(GL_COLOR_BUFFER_BIT);
/* In 3D rendering you'd also clear: GL_DEPTH_BUFFER_BIT | GL_STENCIL_BUFFER_BIT */
/* ── 3. Set the active shader program ──
All subsequent draw calls use this program until you call glUseProgram again. */
glUseProgram(shader_program);
/* ── 4. Bind the VAO ──
This restores all the vertex format state we set up earlier.
The GPU now knows where the vertex data is and how to read it. */
glBindVertexArray(vao);
/* ── 5. Draw ──
glDrawArrays: draw primitives from currently bound VAO.
GL_TRIANGLES: every 3 vertices form one triangle
0: start from vertex index 0
3: draw 3 vertices total (= 1 triangle)
*/
glDrawArrays(GL_TRIANGLES, 0, 3);
/* ── 6. Swap buffers ──
We've been drawing into the BACK buffer.
Swap makes the back buffer visible (front buffer) and
gives us a fresh back buffer to draw the next frame into.
This is double buffering — prevents screen tearing. */
glfwSwapBuffers(window);
}
Putting it all together
Here's the complete program — everything from above assembled into one file you can actually compile and run:
#include "glad/glad.h"
#include <GLFW/glfw3.h>
#include <stdio.h>
#include <stdlib.h>
/* ── Shader source strings ── */
const char *vert_src =
"#version 330 core\n"
"layout (location = 0) in vec3 aPos;\n"
"void main() {\n"
" gl_Position = vec4(aPos, 1.0);\n"
"}\n";
const char *frag_src =
"#version 330 core\n"
"out vec4 FragColor;\n"
"void main() {\n"
" FragColor = vec4(1.0f, 0.5f, 0.1f, 1.0f);\n"
"}\n";
/* ── Triangle geometry ── */
float vertices[] = {
-0.5f, -0.5f, 0.0f,
0.5f, -0.5f, 0.0f,
0.0f, 0.5f, 0.0f,
};
void framebuffer_size_callback(GLFWwindow *w, int width, int height) {
glViewport(0, 0, width, height);
}
GLuint compile_shader(GLenum type, const char *src) {
GLuint s = glCreateShader(type);
glShaderSource(s, 1, &src, NULL);
glCompileShader(s);
GLint ok; glGetShaderiv(s, GL_COMPILE_STATUS, &ok);
if (!ok) {
char log[512]; glGetShaderInfoLog(s, 512, NULL, log);
fprintf(stderr, "Shader error: %s\n", log);
}
return s;
}
int main(void) {
glfwInit();
glfwWindowHint(GLFW_CONTEXT_VERSION_MAJOR, 3);
glfwWindowHint(GLFW_CONTEXT_VERSION_MINOR, 3);
glfwWindowHint(GLFW_OPENGL_PROFILE, GLFW_OPENGL_CORE_PROFILE);
#ifdef __APPLE__
glfwWindowHint(GLFW_OPENGL_FORWARD_COMPAT, GL_TRUE);
#endif
GLFWwindow *window = glfwCreateWindow(800, 600, "Triangle", NULL, NULL);
glfwMakeContextCurrent(window);
glfwSetFramebufferSizeCallback(window, framebuffer_size_callback);
gladLoadGLLoader((GLADloadproc)glfwGetProcAddress);
/* Compile and link shader program */
GLuint vert = compile_shader(GL_VERTEX_SHADER, vert_src);
GLuint frag = compile_shader(GL_FRAGMENT_SHADER, frag_src);
GLuint program = glCreateProgram();
glAttachShader(program, vert);
glAttachShader(program, frag);
glLinkProgram(program);
glDeleteShader(vert);
glDeleteShader(frag);
/* Upload geometry to GPU */
GLuint vao, vbo;
glGenVertexArrays(1, &vao);
glGenBuffers(1, &vbo);
glBindVertexArray(vao);
glBindBuffer(GL_ARRAY_BUFFER, vbo);
glBufferData(GL_ARRAY_BUFFER, sizeof(vertices), vertices, GL_STATIC_DRAW);
glVertexAttribPointer(0, 3, GL_FLOAT, GL_FALSE, 3 * sizeof(float), (void*)0);
glEnableVertexAttribArray(0);
glBindVertexArray(0);
/* Render loop */
while (!glfwWindowShouldClose(window)) {
glfwPollEvents();
if (glfwGetKey(window, GLFW_KEY_ESCAPE) == GLFW_PRESS)
glfwSetWindowShouldClose(window, 1);
glClearColor(0.08f, 0.08f, 0.10f, 1.0f);
glClear(GL_COLOR_BUFFER_BIT);
glUseProgram(program);
glBindVertexArray(vao);
glDrawArrays(GL_TRIANGLES, 0, 3);
glfwSwapBuffers(window);
}
glDeleteVertexArrays(1, &vao);
glDeleteBuffers(1, &vbo);
glDeleteProgram(program);
glfwDestroyWindow(window);
glfwTerminate();
return 0;
}
Debugging OpenGL errors
OpenGL's error handling is historically terrible. Functions don't return error codes — errors accumulate in a queue that you have to poll manually with glGetError(). It's asynchronous and easy to ignore.
/* Basic error checking — call after any GL operation you're suspicious of */
void check_gl_error(const char *context) {
GLenum err;
while ((err = glGetError()) != GL_NO_ERROR) {
const char *desc;
switch (err) {
case GL_INVALID_ENUM: desc = "GL_INVALID_ENUM"; break;
case GL_INVALID_VALUE: desc = "GL_INVALID_VALUE"; break;
case GL_INVALID_OPERATION: desc = "GL_INVALID_OPERATION"; break;
case GL_OUT_OF_MEMORY: desc = "GL_OUT_OF_MEMORY"; break;
default: desc = "unknown error"; break;
}
fprintf(stderr, "GL error in [%s]: 0x%x (%s)\n", context, err, desc);
}
}
/* Usage: */
glBindVertexArray(vao);
check_gl_error("after glBindVertexArray");
glDrawArrays(GL_TRIANGLES, 0, 3);
check_gl_error("after glDrawArrays");
Modern OpenGL (4.3+) introduced Debug Output — a much better system that calls a callback function whenever an error occurs, with the error details, severity, and even which GL call caused it. If your hardware supports it (most modern GPUs do), use it:
/* OpenGL 4.3+ debug callback — set this up right after context creation */
void APIENTRY gl_debug_callback(
GLenum source, GLenum type, GLuint id, GLenum severity,
GLsizei length, const char *message, const void *userParam)
{
/* Filter out low-importance notifications */
if (severity == GL_DEBUG_SEVERITY_NOTIFICATION) return;
fprintf(stderr, "GL Debug [severity=%s]: %s\n",
severity == GL_DEBUG_SEVERITY_HIGH ? "HIGH" :
severity == GL_DEBUG_SEVERITY_MEDIUM ? "MEDIUM" : "LOW",
message);
}
/* In your setup code (after gladLoadGLLoader): */
glEnable(GL_DEBUG_OUTPUT);
glEnable(GL_DEBUG_OUTPUT_SYNCHRONOUS); /* callback fires on the same thread = better stack traces */
glDebugMessageCallback(gl_debug_callback, NULL);
glEnableVertexAttribArray — triangle is invisible. (2) Wrong stride in glVertexAttribPointer — triangle is garbage/stretched. (3) Calling glVertexAttribPointer before binding the VAO — attribute state saved to wrong object. (4) Shader compile error not checked — draws nothing, no idea why. (5) Forgetting glfwPollEvents() — window appears frozen.
What comes next
The orange triangle on a dark background is anticlimactic but it's genuinely the hardest part — not because it's complex, but because there's so much one-time setup before anything appears. From here, each new feature is much more incremental:
- Uniforms — pass values from your CPU code into shaders at draw time. This is how you animate a triangle (pass a time value, offset position in the shader) or tint its color dynamically.
glUniform1f,glUniformMatrix4fv, etc. - Multiple attributes — add per-vertex color by adding RGB data to your
vertices[]array and a secondglVertexAttribPointercall. The rasterizer will interpolate colors across the triangle face automatically. - Textures — load an image, upload it to the GPU, sample it in the fragment shader using texture coordinates (UV values per vertex). Adds two new objects:
GL_TEXTURE_2Dand a sampler uniform. - 3D transformation — add a
mat4uniform to the vertex shader and multiplygl_Position = projection * view * model * vec4(aPos, 1.0). Use a math library likecglmorglmto build the matrices on the CPU side. - Depth testing — enable
GL_DEPTH_TESTand clearGL_DEPTH_BUFFER_BITeach frame. Now closer geometry automatically occludes farther geometry without you having to sort draw calls. - Index buffers (EBO) — instead of repeating vertex data for shared vertices (a quad needs 4 vertices, not 6), use an index buffer to reference vertices by index. Add a
GL_ELEMENT_ARRAY_BUFFERand switch fromglDrawArraystoglDrawElements.
Each of these adds one or two new concepts but reuses everything you've already learned. The pipeline doesn't change — you just give it more to work with. The hard mental model work is done the moment you understand how that first triangle gets to the screen.
One parting observation: the reason Vulkan (and Metal, and DX12) feel so much harder than OpenGL is not because they do anything fundamentally different — it's that they make explicit all the things OpenGL hides from you: memory management on the GPU, render pass structure, synchronisation between CPU and GPU work, pipeline state objects. OpenGL's "magic" is Vulkan's explicit setup. Understanding OpenGL deeply makes Vulkan feel like "oh, this is just exposing what I already knew was happening."
So draw the triangle. Then change its color. Then animate it. Understand each piece before moving to the next. The whole world of real-time graphics opens up from exactly this starting point.