How Were Computer Graphics Invented?

https://www.youtube.com/watch?v=KILGR41vIeE

无作者

Today’s 3D graphics look better than ever, but at their core,these visuals haven’t changed all that much in over 50 years.The inventions made in the mid-20th century have largely remained, as the base components of 3D worlds to this day.People have been using these building blocks ever since,but how did they make the blocks themselves?Prior to the 1950s, computers were used for solving complex mathematical, problems, processing government data, and breaking military codes.For these purposes, imagery wasn’t essential.But with the start of the cold war,the first real need for data visualization arose.The U.S.Naval Research Lab wondered whether it, possible to create a flight simulator for training bomber crews.The idea was to simply update simulated instruments based on pilot inputs.They reached out to MIT, who concluded that the project was feasible.Inspired by ENIAC, they built a digital computer,the first in the world to perform parallel bit calculations.

In the end, the computer became the basis for the air defense system,displaying aircraft and missile positions on CRT screens.This was the first instance in history where data was processed, specifically for the purpose of being visualized.Okay, this is cool and all, but it’s far from 3D.The thing is, uncovering the very first instance of someone adding, a third coordinate to a point inside a of computer is surprisingly difficult.The earliest example of a 3D visual I could find is this 1961, Swedish animation of a planned highway.There is clear perspective shift in the video,so some form of 3D calculations were in place,but it’s unclear what the math actually looked like.The first clear example we have is the software Sketchpad, developed by Ivan Sutherland and Timothy Johnson in 1963.They actually used 4 coordinates.Okay, hold on, I’m getting ahead of myself.To best understand why they used what they did,we need to understand what they were trying to solve in the first place.

The first obvious question regarding 3D, graphics is; how should it be visualized on a 2D screen?At the time they used oscilloscopes with 1024, x 1024 discrete points, aka pixels.Let’s make the x and y coordinate of our 3D world, the screen itself and make the center the origin.So, if we have a 3D point at coords (24,12,4),we can render it 20 to the right, and 12 up from the center.For now, we ignore the z coordinate.This is called an orthogonal projection; there is no perspective.Another point at (24,12,12), will render into the same pixel, despite being further away.To get perspective, we can simply divide the x and y by the Z coordinate.So, the two points become (6,3) and (2,1) respectively.The more the depth increases, the closer to the center the points will move.This is basically a one-point perspective that you may have used in art class.Now, what if we have a, shape described by multiple points and we want to transform it?We could grab each of the points and do the math one by one; add.

the rotation, scale it, move it, but mathematics, already has a solution for doing all of it in one go called the matrix.We can make new points by applying matrices to existing ones.For example, we can double the distance of this 2D, point from the origin using this matrix.It takes 2 of the X plus 0 of the Y value to construct a new X.The Y will become 0 times X plus 2 times Y.We can rotate the point around the origin too.Let’s rotate it by 90 degrees clockwise.For that X needs to become Y and Y needs to become negative X.But if we want to move the point, we run into a problem.The matrix doesn’t allow us to add constants.In its current form it can only reference the X and Y values,so let’s add another coordinate called the scale and make it 1.Now if want to move the point say 5 right and 3 up,we can simply make X be itself plus 5 of the constant.Y becomes 1 times Y plus 3 times the constant and the scale is kept, the same. Now we can pack all three of these transformations.

into a single matrix and apply it to every point of a shape.But on top of this, our new coordinate also helps with projection.We can use it to save the Z value, for when we need to determine what is closer to the camera.It also incorporates the perspective math into the final matrix.Don’t sweat the details of how exactly the extra coordinate achieves this,I had multiple aneurysms trying to make sense of it.The bottom line is that basically everything is in matrices, now, which computers can work very efficiently with. The lines, themselves are actually only drawn between the projected points, at the very end to create simple wireframe renders.And so, the first 3D models were born, including the first human created by William Fetter.It was a little cursed with every line rendered.But thankfully, Lawrence Roberts figured out that lines can be hidden, based on whether they faced the camera and convex objects could even block other, objects’ lines, like in this animation created by Edward Zajac.

Sketchpad would end up having a massive influence on modern software,but this was only the start of Sutherland’s career,he would go on to work on even more influential projects at a university, that would become known as the birthplace of computer graphics, and where the rest of our rediscovery takes place.In 1965 David Evans, started his computer science department at the University of Utah.After a bit of convincing, he managed to get Sutherland on board.The two even started their own company and attracted a lot of brilliant minds, to join their class.Evan’s students included Edwin Catmull, Bui Tuong Phong, Jim Blinn,and Henri Gouraud.If you know your way around CG, then some of these names should sound, very familiar to you. With CG, still very much in its infancy, the group had a lot of problems to tackle.3D models were still just points and lines, so they decided to fill, the spaces between them with so-called faces or polygons.It’s kind of like that magnetic game.

You were missing out if you didn’t have this as a kid by the way.But how do you transform the 3D faces onto a flat screen?How do you properly render them on top each other?And how do you light them? Evans, and three of his students found solutions to all these problems.Their answer to the first question was to matrix transform everything, so that the viewpoint becomes the world origin and also rotate the world, so that it lines up with the camera’s axis.Just like with Sketchpad, the camera’s forward vector becomes the Z axis, and everything in the world is now relative to the camera.Today we call this eye, or camera, or view space.This makes it easier to perform the projection. We can imagine, the screen as a rectangle placed in front of the camera.To figure out what should be rendered to a pixel, lets draw, a straight line from the view point through that pixel into the scene.The distance we want this ray to travel effectively determines the far clipping, point of the camera.

If it hits nothing, we take the background color.If it finds one polygon, we take that polygon’s color.If it intersects multiple polygons, we calculate each intersection’s, distance along the line and take the closest one.This was an early form of the depth- or Z-buffer, and solves the second problem as well.As for the last question, they figured that using triangles, would be the simplest.To tell how bright one should appear, they, first calculated the cross product of two of its edges.After normalizing, the resulting vector is called the normal,which points perpendicularly away from the face.To avoid dealing with shadows,they considered the light source to be the view point itself,so the following calculations were all done in view space.To get the direction to the light they could simply subtract, the position of the view point from the face’s position.After normalizing this as well, we can take the dot product of the two vectors.

The result will fall between -1 and 1 depending on the angle between them.If they point in the same direction the result is 1 and as the angle, increases, the dot product lowers reaching 0 at 90 degrees.Going past 90, the value behaves the same but goes negative.We can clamp the value at 0 to avoid the negative numbers,but triangles facing away from the light also face away from the camera,so they can be just disregarded by the renderer.This is called backface culling and it’s why, if you clip through the ground, in a game, that you can see through surfaces.To finalize the lighting calculations, we can simply multiply each triangle’s, brightness by their dot product. This was a great start, and allowed for the rendering of clean and recognizable 3D models.However, complex renders took minutes to finish.Thankfully, a good number of optimizations were found.John Warnock split the image into smaller sections based on simplicity, and rendered each at once instead of going pixel by pixel.

Gary Watkins streamlined the whole rendering process, by minimizing recalculations, memory accesses, and fitting the math into simple loopable functions.This was so effective that they could render scenes in real time at 30 fps.But rendering smooth surfaces was still impossible, unless you used a lot of triangles, which is impractical and caused artifacts, This is where Henri Gouraud comes into the picture.His idea was to somehow blend the shading across the faces.He started by adding normals at the vertices by averaging the face normals.He then performed the same calculations to get the light intensity, at those points. The next step takes place after projection,which is a number of linear interpolations.First, we interpolate across the edges.Then if we select a row of pixels that crosses the face,we will hit at least 2 edges that we can interpolate from.Once this is done for every row and every face, we have Gouraud shading.Gouraud tested it on a number of 3D models including his wife’s face.

and although the algorithm has evolved over time, the base principle remains, the same and is used to this day, mainly for viewport rendering.In 1972, Sutherland challenged his students to model something iconic.They chose a Volkswagen Beetle, since Ivan’s wife had one.They drew and measured the wireframe on the car itself and inputted the numbers.This is the actual 3D model.Topology could be better, but hey, they had to type each of the vertices, in one by one and with Gouraud’s new shading, it doesn’t look half bad.Although one of the students, Bui Tuong Phong,pointed out that the shading still had discontinuities.The brightest point on a curved surface could, and most often will fall on the face, which Gouraud’s technique will miss.So, he proposed a new way to light the object.Instead of interpolating the brightness of the vertices across the face,Phong interpolated the normals themselves.And that’s all there really is to it.Instead of blending the resulting light, Phong blends.

the surface direction and then calculates the light.While this increases computation time, it creates a proper,smooth appearance even on low poly meshes.But what Phong was really interested in was physically accurate lighting,which meant the introduction of specular reflection.This was calculated by bouncing the light off of the object.The appropriate angle is easily calculated from light’s direction and the normal.Then its direction is checked against the camera’s location.The closer the reflected, light points to the viewer the brighter that spot will appear.This value is added on top of the normal based shading to create the final result.Phong also generalized his algorithms to allow objects to be lit, from any direction. The fundamentals of 3D modelling were now in place,but the University of Utah was not done innovating.Edwin Catmull realized that pictures can be mapped onto objects, by using another coordinate system; the UV space.

This is a 2-dimensional layout of the 3D model’s flattened surface.You can also think of it, as each vertex of the mesh being assigned a u and v coordinate,which determines what part of a texture should be applied there.A year later, the most famous 3D model was created; the Utah teapot.Martin Newell needed something familiar for his work.His wife, Sandra, suggested modelling their tea set.Newell used Bézier curves, which marked one of the first instances, of a solid 3D model being made with mathematical formulas.The teapot became the standard reference test model for 3D rendering.Newell himself used it for his and Jim Blinn’s reflection algorithm,which utilized a new spherical texture called the environment map.For each surface point of a mesh the camera direction and surface, normal can be used to calculate the reflection direction.This is converted, into spherical coordinates and used to sample the environment map.The values are added on top of the surface shading.

Blinn also wanted to give objects, a more detailed surface, so he came up with bump mapping.This uses a separate texture called a heightmap.Before light calculation, the heightmap alters the surface normals.Note that for areas that contain the same heightmap value, the normal stays, the same.Areas with changing values is where the normal is tilted.Computing the shading with the new normals give the appearance, of an angulating surface, even though the physical mesh remains the same.A problem that no one managed to solve at Utah, until now is aliasing; an ugly side effect of using pixels.Graphics are aliased if they appear jagged or pixelated.This was most prominent along edges of object silhouettes.Pixels can only have a singular color, even if they happen to cover an area with multiple values.Franklin Crow realized that instead of sampling, just one point for a pixel, they could do 4 and average them.We can subdivide further of course to get even more accurate colors,

but even just one subdivision considerably increases computing costs.Especially, since Crow subdivided the full screen.These days we refer to this as supersampling.It’s basically rendering the scene with quadruple the pixels.A number of more efficient anti-aliasing solutions, have been developed since, all with their pros and cons.But with that we have reached the last of the legacy that Utah’s school, of computing left behind without which, 3D graphics wouldn’t be today.They of course didn’t invent everything; it would take another decade of progress, to get to a point where computer graphics could be commercialized.But that’s a topic for another day.