analysis
CGI: Why Images Are No Longer Just Shot — They’re Assembled
3D scenes, virtual light, compositing, reconstruction, and generative video as a new production logic of the frame

The camera increasingly records only part of the image. The rest can be modeled, calculated, reconstructed, composited, or synthesized—and assembly becomes the new center of visual production.
Editorial question
If an image can be captured, modeled, reconstructed from photographs, assembled from layers, or synthesized by a model, what does it mean today to “shoot a frame”?
Short answer
CGI turned the image from the output of one camera into a production system. Jurassic Park demonstrated the credibility of hybridity; Toy Story and RenderMan showed the possibility of an entire computable world; OpenUSD made the scene collaborative infrastructure; StageCraft returned digital environments to the set; NeRF and 3D Gaussian Splatting turned photography into data for new viewpoints; generative video made motion possible without exposing an explicit 3D scene.
The main change is therefore not that “everything became digital.” It is that the image became rebuildable. The author’s question shifts from “what was in front of the camera?” to “which dependencies create the final picture, and who controls them?”
Thesis
CGI changed visual culture not because it made a “fake” image indistinguishable from a real one. Its deeper shift is that the frame stopped being only a record of an event and became a designed scene that can be rebuilt—from geometry and light to camera, compositing, reconstruction, and generative layers.
The camera does not disappear. It becomes one input into a wider system. A contemporary frame can contain captured, modeled, calculated, reconstructed, and synthesized imagery at the same time. This hybridity—not a simple opposition between “real” and “digital”—defines the new literacy of image production.
The camera increasingly records only part of the image. The rest can be modeled, calculated, reconstructed, composited, or synthesized—and assembly becomes the new center of visual production.
Evidence and cases
The dinosaur the camera never photographed
In Jurassic Park, the audience had to believe in a creature that did not exist in front of the lens. Not an abstract “digital dinosaur,” but a heavy animal that occupies space, catches light, casts a shadow, and moves beside physical sets. The Academy still describes the film as a benchmark in CGI and as part of the transition from stop-motion photography to CG animation. 1
This is where the central change becomes visible. Computer graphics became culturally important not merely when they could show the impossible, but when viewers stopped noticing the production boundaries inside a single frame. Animatronics, camera optics, lighting, digital geometry, textures, and compositing began to operate as one visual system.
So the question of this article is not “how realistic has CGI become?” It is more difficult: what happens to the photographic frame when its final appearance no longer has to exist completely on set, in one file, or even in one type of image?
The frame stops being only a recording and becomes a scene
Photography and cinema traditionally linked the image to an event in front of the camera: light from the world passes through optics and leaves a record. Digital production does not erase that relationship, but it makes it only one option. Light can be calculated, an object modeled, a material parameterized, motion simulated, a background reconstructed from photographs, and part of a shot synthesized by a model.
The result is a fundamentally different object: not only a frame, but a scene that can be rebuilt. Camera position, light intensity, metal roughness, glass transparency, fabric speed, and atmospheric depth can be changed after an initial decision has already been made. The image becomes versionable.
Terminology matters. CGI means computer-generated image elements. VFX is a broader production practice that may include CGI, compositing, and other methods. Virtual production connects digital scenes with physical filmmaking in real time. Generative video may become another input into the same pipeline, but it is not a synonym for CGI. Contemporary frames often combine several of these regimes at once.
Jurassic Park: credibility is born from hybridity
Jurassic Park is most useful not as a story of digital victory over analog craft, but as a lesson in hybridity. The film won the Academy Award for Visual Effects, and the Academy emphasizes its transitional role between stop-motion and CG animation. 1
The point is not that computers replaced physical effects. Credibility emerged from coordination between sources. A digital creature had to obey the cinematographic world already in place: light direction, perspective, camera motion, scale, atmosphere, and editing rhythm. If one layer betrayed itself, the dinosaur lost weight.
That remains one of the core laws of convincing CGI: viewers do not believe a polygon or a texture; they believe relationships. An object becomes real to the eye when it behaves as if it belongs to the same world as the actor, the floor, the rain, and the lens.
Toy Story and RenderMan: the image becomes a production discipline
Two years after Jurassic Park came Toy Story. Pixar calls it the world’s first computer-animated feature film, while the history of RenderMan connects the rendering system to the production of that fully digital feature world. 3 [source]
Here CGI makes the next move. It no longer inserts a few impossible objects into filmed reality; it carries the entire frame. Character, room, material, highlight, depth, camera, and motion exist inside one computable environment. This requires not one spectacular effect but discipline: the scene must survive thousands of frames, viewpoints, and revisions.
That is why RenderMan matters historically not simply as a software brand, but as a sign of a production shift. Rendering becomes repeatable. Light can be recalculated, a material changed, an object replaced without destroying the entire scene. The frame begins to behave less like a unique imprint and more like a state of a complex system.
Material, light, and camera: reality becomes parameters
In a 3D scene, a surface receives behavior. Metal is not defined by the word “metal,” but by its relationship to light: reflection, roughness, microstructure, color, transparency, and other properties. The same is true of fabric, leather, glass, stone, or skin. Rendering does not merely draw a material; it calculates how that material should respond to its environment.
For a photographer, this almost reverses the familiar order. In a physical studio, light is arranged around an existing object. In a digital scene, the light, the object, and the surface can all change together. Product imagery makes this obvious: one digital twin of a watch, piece of jewelry, or bottle can exist in many angles and lighting setups without a repeated physical shoot.
Freedom, however, increases the demand for observation. The more parameters can be controlled, the easier it is to create an image that is perfectly clean and completely lifeless. Realism comes not from maximum polish but from convincing constraints: weight, imperfection, highlight scale, surface variation, fabric inertia, and believable optics.
Compositing: the final frame is usually more complicated than one render
The image rarely ends at the render. Compositing brings together live-action footage, CG passes, mattes, depth, haze, reflections, color, grain, and local corrections. Foundry describes Nuke as a node-based environment for digital compositing, where operations form a network of dependent elements. 4
Node-based thinking explains why “assembled” is more precise than “drawn.” The artist sees not a pile of anonymous layers but a chain of causes: where a matte came from, where color changed, which pass controls a highlight, how depth affects atmosphere. The frame becomes a map of decisions.
This matters to visual culture because final credibility often appears precisely at the boundary between sources. CG may be too clean; footage too noisy; a reflection physically correct but visually weak. Compositing does not merely glue elements together. It brings them into a shared regime of reality. 4
OpenUSD: the image is no longer a file
Once a digital scene becomes large, the problem is no longer the beauty of a single object. The problem is how many people and applications can work on one reality without destroying each other’s work. OpenUSD grew from Pixar’s production needs and is described as a platform for collaboratively constructing large 3D scenes containing geometry, shading, lighting, animation, and variants. 7
This is one of the most underestimated shifts. A contemporary digital frame is not necessarily one file opened by one artist. It may be composed from many assets and versions: a character from one department, clothing from another, environment from a third, lighting from a fourth. The scene composes them without flattening everything into an irreversible final state.
For brands and architecture, this logic is especially powerful. One well-prepared digital object can live in a campaign, film, configurator, interactive scene, and future visualization. The image begins to inherit the properties of infrastructure: it can be not only viewed, but continued.
StageCraft: CGI returns to the set
Virtual production changes the familiar sequence of “shoot first, add later.” ILM describes StageCraft as an end-to-end virtual production system; on the first season of The Mandalorian, actors performed inside a large LED environment where digital backgrounds surrounded the set and could react to the camera in real time. 9
This is a crucial turn. The digital background no longer waits for post-production—it influences physical photography immediately. Light from LED panels reaches costume and skin, the cinematographer sees composition, the director can alter the environment, and actors receive a space rather than an abstract green screen. Epic describes virtual production more broadly as the real-time joining of digital and physical worlds. [source]
The boundary between pre-production, production, and post-production becomes less linear. Decisions that were once postponed now have to be made earlier: world scale, light direction, perspective, and camera behavior. CGI returns to the set, but it brings a new cost—the need to design the frame sooner.
NeRF and 3D Gaussian Splatting: the camera becomes a scene sensor
For a long time, digital scenes were mainly built explicitly: geometry was modeled, materials assigned, lights placed. Neural rendering proposed another route. NeRF showed that a continuous scene representation could be optimized from a set of images with known camera poses and then used to synthesize novel viewpoints. 14
3D Gaussian Splatting makes a different trade-off between scene representation, quality, and speed. Its authors represent scenes with 3D Gaussians and report real-time novel-view rendering on their evaluation setups. [source]
For photographers, the change is almost philosophical. A set of photographs stops being only a set of final images. It becomes a measurement of space—raw material from which a new scene and a new virtual camera can emerge. Photography becomes data collection for an image that does not yet exist.
Generative video: when even an explicit 3D scene is no longer mandatory
Generative video models add another kind of assembly. In classical CGI, an artist explicitly creates or imports geometry, materials, lighting, and animation. A generative model can synthesize appearance and motion from text, images, video, and references without exposing an explicit 3D scene to the author.
Runway described Gen-4 in terms of consistency of characters, objects, and locations across scenes, and Gen-4.5 in terms of improvements in motion quality, prompt adherence, and temporal consistency. These are vendor claims and should be read as such, but they reveal the direction of the market: value is shifting from a single spectacular generation toward repeatable, controllable worlds. 16 [source]
Interpretation
Generative video does not abolish CGI. It makes the next stage of the same cultural change visible: the viewer cares about a coherent frame, while the producer cares about control over its internal dependencies. If 3D offered control through an explicit scene, a generative model tries to offer it through references, context, and instructions. The less of the internal scene the author can see, the more important the question becomes: where, exactly, does direction now live?
For VANSMITHLAB this is also a question of trust. An image may be photographically convincing without one photographic event beneath it. Provenance, rights, consent, reconstruction labels, and the distinction between document and illustration therefore become part of visual literacy, not merely technical notes.
Not a fake image, but designed reality
CGI is often discussed through a simple opposition: real or artificial. That opposition no longer describes the contemporary frame very well. Jurassic Park was a hybrid of physical and digital production. Toy Story existed as a fully computable scene. StageCraft returned digital environments to the physical set. NeRF and Gaussian Splatting turn photographs into material for a new camera. Generative video can create motion without a conventional 3D scene.
The common denominator is not “fakeness.” It is designability. The image becomes a system in which form, light, material, motion, viewpoint, time, and provenance can be controlled separately. The camera remains one of the most important tools, but it is increasingly one input among several rather than the only source of the frame.
That is why the dinosaur from the opening still matters. It was convincing not because a computer learned to impersonate reality, but because different kinds of reality were assembled so precisely that the viewer stopped seeing the seam. Today that principle extends far beyond visual effects. We increasingly do not simply shoot an image—we design the conditions in which it can exist.
Visual analysis
The article's visual logic can be tested through «The dinosaur the camera never photographed», «The frame stops being only a recording and becomes a scene», «Jurassic Park: credibility is born from hybridity», «Toy Story and RenderMan: the image becomes a production discipline», «Material, light, and camera: reality becomes parameters», «Compositing: the final frame is usually more complicated than one render». Compare these signs within one object or sequence rather than judging them as an isolated style.
Practical implication
Practical use begins by testing the thesis against a specific object, frame, or production decision. For this subject, the working criterion is: 3D scenes, virtual light, compositing, reconstruction, and generative video as a new production logic of the frame
Limitations
“CGI: Why Images Are No Longer Just Shot — They’re Assembled” is limited to its stated subject: 3d scenes, virtual light, compositing, reconstruction, and generative video as a new production logic of the frame. The analysis does not replace examination of a specific original, project, or production context.
Sources
The open sources for “CGI: Why Images Are No Longer Just Shot — They’re Assembled” are listed in the page panel and define the boundary of its verifiable facts.