ANNIE BURKE
RESEARCH / FLORIDA INTERNATIONAL UNIVERSITY

Work in progress.

Inferential discovery in generative workflows:
how abundance affects the designer’s relationship to work in progress.

2024 — 2026
01 / CURRENT INSTRUMENT

Inference Resolution Map.

Previous iterations ↘
Earlier placements Revised X
Inference Resolution Map, working comparisonPeach-to-yellow circles show earlier recorded placements. Bright diamonds show revised provisional X placements at retained earlier Y values. Select a mark to inspect records. No coordinates can be edited.
More opportunity for inferenceLess opportunity for inference →
X / OPPORTUNITY FOR INFERENCEY / EARLIER EDITABILITY VALUES RETAINED
Working placements · assessment notes through 22 September 2026 Read-only catalogue
Placement methodology +

The equation measures automation on a linear scale, providing a basis for discussing a spectrum from automation to agency.

CURRENT WORKING EQUATION / X
X = 15 − 3∑Oᵢ

Oᵢ is the assessed opportunity for inference in each of five aspects: subject, composition, form, appearance, and detail. Each is rated from 0 to 1, in provisional quarter steps. Higher opportunity places the operation farther left; lower opportunity places it farther right.

The sum ranges from 0 to 5. Multiplying by 3 maps it onto 0–15; subtracting from 15 preserves the map’s orientation. Each aspect contributes equally, with a maximum contribution of 3 coordinate units.

Worked calculation +

ComfyUI text-to-image example: a scoped assessment, not a rating of every ComfyUI operation.

SubjectCompositionFormAppearanceDetail
0.751.000.750.501.00

∑Oᵢ = 4.00
X = 15 − (3 × 4.00) = 3.00

A change of 0.25 in one aspect changes X by 0.75. For example, lowering detail opportunity from 1.00 to 0.75 produces X = 3.75. This sensitivity calculation shows the consequence of a disputed rating; it does not change the plotted assessment.

Different profiles can have the same total. The five ratings and their rationales therefore remain available beside the coordinate.

Earlier equations +

Initial placements

The initial survey used shared axis descriptions and reference examples to guide coordinate judgments. A calculation for every original decimal coordinate has not been recovered. Those placements are part of the instrument’s history.

Four-component proposals

The archived Cell Math discussion contains successive proposals with different targets:

Xautomation = (10 / 8)(T + G + C + R)

Task delegation, generative resolution, decision closure, and run/continuation autonomy were each rated 0–2.

Xspecification = (10 / 8)(D + R + C + U)

A later proposal used direct input, representation, constraint control, and retention of user decisions, again rated 0–2. Both convert a maximum sum of 8 to a 0–10 scale. With integer component ratings, possible totals are separated by 1.25. These are historical proposals, not the equation used for the revised points.

Five-aspect specification score · September 20

Xspecification = ∑Sᵢ   ·   Sᵢ ∈ {0, 1, 2, 3}

Subject, composition, form, appearance, and detail were scored by how they were supplied: unspecified, described in words, constrained by a guide or source, or explicitly constructed. Their sum gave a 0–15 specification score; the accompanying notes proposed 15 − X as opportunity for inference.

Five-aspect opportunity assessments · September 21 onward

X = 15 − 3∑Oᵢ

The current worked examples assess what supplied inputs leave open for generative AI to resolve within a specific operation. They retain the five aspects while documenting the scenario, source evidence, rating rationale, and sensitivity. This is a change in assessment conventions, not merely a conversion of old scores. The earlier and revised points should not be read as measured changes in software.

Judgment, consistency, and calibration +

The wine-rating analogy motivates a common descriptive reference: people can inspect a rating without sharing a preference. Here, the coordinate describes an operation’s production conditions; it is not a quality or taste score.

The arithmetic is repeatable once the five ratings are fixed. The research task is to make those ratings defensible: specify the operation and inputs, document evidence for each aspect, expose uncertainty, and compare independently assigned ratings. Independent agreement and calibration have not yet been established for this working rubric.

Equal weights and equally spaced quarter steps are provisional assumptions. They allow an explicit calculation, but do not establish that the aspects are equally consequential or that perceived differences are equally spaced. A decimal coordinate is not a percentage of AI contribution or a direct count of inferences.

This preview contains 387 operation records and 81 revised X examples. Earlier Y values are retained for comparison and have not been reassessed. Records without Y remain in the catalogue. Original source coordinates remain unchanged.

Operation catalogue 387 records+

Search above to narrow results. Open a record to inspect its assessment.

Previous iterations+
29 AUGUST 2026

Photoshop operation comparison

AUGUST 2026 / PROTOTYPE

Earlier map

02 / INQUIRY

AutomationAgency

Inferential Discovery and Design Adaptation in Generative Workflows

The inquiry is located at the designer’s point of discovery.

This dissertation examines how the designer responds, and how an abundance of new information provided by generative workflows affects their relationship to work in progress.

A discovery can change how the designer understands the work and what its continuation might mean. Following these responses makes it possible to examine how that relationship develops, while further discovery remains possible.

The Inference Resolution Map describes production conditions that help situate these responses. It supports comparison across operations and provides context for studying the designer’s response and its significance for continuation.

RESEARCH QUESTION

How does a designer’s relationship to their visual work in progress shift when inferential discovery can recur across a generative workflow?

Research proposal 11 September 2026+Download PDF ↓

Dated proposal. The map and its placement criteria continue to develop.

Next steps

The abstract, introduction, and Section 1 are drafted, and Section 2 is mapped out. The next phase develops the instrument and brings it into the practice-based inquiry that will inform Sections 3–5.

Expand the assessments

Continue developing the map’s operation-level points, with documented inputs, placement criteria, evidence, and rationales. Compare cases and examine how uncertain ratings affect their positions.

Investigate Y

Examine whether and how persistence, continuance, editability, and aesthetic considerations relate to one another—and what, if any, basis they offer for the Y axis. Their relationship and suitability for a shared axis remain open questions.

Bring the inquiry into practice

Use my own work to examine the designer’s response at the point of discovery and what it means for continuation. Develop Sections 3–5 through the relationship between these responses, the map’s production conditions, and the evidence gathered in practice.

03 / 2024 — 2025

Related research

2025

Modalities & spatial structures

Multimodal Spatial Intelligence
+

Earlier work examined how different input modalities supply partial spatial information, how generative systems resolve that information, and how people interpret the resulting structures.

This provides context for the ongoing inquiry into what is supplied, what is inferred, and what a designer encounters in the work.

Read 2025 paper ↗
2024–25

Teaching, tools, and individual learning

+

While teaching architecture in Miami, I explored how people could learn to use tools through experiences tailored to them. My early research considered visualization and gamification as ways of making learning interactive, with climate education as an initial context. The emphasis subsequently moved toward personalized education and how learning could respond to individual needs.

As generative practice became a stronger focus, my attention shifted toward what designers encounter while working with these tools. The Inference Resolution Map sits within this longer inquiry into people’s relationships with tools; the current dissertation focuses on the designer’s response to inferential discovery.

Research sequence +
OCTOBER–NOVEMBER 2024

Visualization and interactive learning

Literature review and research-paper drafts examined 3D visualization, gamification, and AI-supported personalization in climate education.

DECEMBER 2024

Architectural education and AI

A Digital Futures paper considered prospective changes to architectural education, curricula, and relationships between educators and students.

MARCH 2025

Personalized learning

Personalized Learning in Higher Education: AI-Driven Analytics and Competency-Based Models as Catalysts for Reform placed adaptive learning and individual needs at the center of the inquiry.

Personal reading & sources

Personal reading+

Listening is an important part of how I learn and develop my work. My reading and listening span technology, creativity, perception, culture, and lived experience.

A selection from my reading and listening, including Algorithms to Live By, Visual Thinking, Unmasking AI, music biographies, and cultural history
Browse titles 78 titles+

78 titles

  • Feeling BeautyG. Gabrielle Starr
  • RangeDavid Epstein
  • The Nvidia WayTae Kim
  • PeakAnders Ericsson and Robert Pool
  • A Paradise Built in HellRebecca Solnit
  • Algorithms to Live ByBrian Christian and Tom Griffiths
  • The Anxious GenerationJonathan Haidt
  • Building a StoryBrandDonald Miller
  • The Thinking MachineStephen Witt
  • The Infinity MachineSebastian Mallaby
  • Worldwide Evil and Misery – The Legacy of the 13 Satanic BloodlinesRobin de Ruiter and Fritz Springmeier
  • Life on EarthDavid Attenborough
  • Breaking the Habit of Being YourselfJoe Dispenza
  • Blue MindWallace J. Nichols
  • OceanDavid Attenborough and Colin Butfield
  • Africa Is Not a CountryDipo Faloyin
  • Exactly What to SayPhil M. Jones
  • GoogooshGoogoosh and Tara Dehlavi
  • QuantumManjit Kumar
  • Visual ThinkingTemple Grandin
  • The Let Them TheoryMel Robbins
  • Only God Can Judge MeJeff Pearlman
  • It Was All a DreamJustin Tinsley
  • Tupac ShakurStaci Robinson
  • Unmasking AIJoy Buolamwini
  • Energy Healing for AnimalsJoan Ranquet
  • How the Word Is PassedClint Smith
  • CasteIsabel Wilkerson
  • Re-RegulatedAnna Runkle
  • Miracle and Wonder (Special Edition)Malcolm Gladwell and Bruce Headlam
  • David and GoliathMalcolm Gladwell
  • Claim Your PowerMastin Kipp
  • Your Debt PlanRachel Rodgers
  • AbundanceEzra Klein and Derek Thompson
  • WordslutAmanda Montell
  • Reclaim Your Nervous SystemMastin Kipp
  • The Mountain Is YouBrianna Wiest
  • The Creative ActRick Rubin
  • BlinkMalcolm Gladwell
  • I Hate the Ivy LeagueMalcolm Gladwell
  • Data FeminismCatherine D’Ignazio and Lauren F. Klein
  • OutliersMalcolm Gladwell
  • SapiensYuval Noah Harari
  • ShowboatRoland Lazenby
  • Revenge of the Tipping PointMalcolm Gladwell
  • NexusYuval Noah Harari
  • The Tipping PointMalcolm Gladwell
  • What the Dog SawMalcolm Gladwell
  • Finding MeViola Davis
  • SupercommunicatorsCharles Duhigg
  • DuneFrank Herbert
  • The Deep Learning RevolutionTerrence J. Sejnowski
  • Weapons of Math DestructionCathy O’Neil
  • Invisible WomenCaroline Criado Perez
  • Stop Living on AutopilotAntonio Neves
  • It’s All in Your HeadRuss
  • The Broken LadderKeith Payne
  • Evolve Your BrainJoe Dispenza
  • The Four Horsemen
  • We Should All Be MillionairesRachel Rodgers
  • Steve JobsWalter Isaacson
  • Elon MuskAshlee Vance
  • Healing Your Attachment WoundsDiane Poole Heller
  • The AlchemistPaulo Coelho
  • The Body Keeps the ScoreBessel van der Kolk
  • The Autobiography of Malcolm XMalcolm X and Alex Haley
  • Shoe DogPhil Knight
  • WillWill Smith and Mark Manson
  • Born a CrimeTrevor Noah
  • JAY-ZMichael Eric Dyson
  • Dark PsychologyValerie Glossner
  • MetahumanDeepak Chopra
  • Atomic HabitsJames Clear
  • AttachedAmir Levine and Rachel Heller
  • Trillion Dollar Coach
  • The Spark and the GrindErik Wahl
  • Summary: 12 Rules for LifeEpicread
  • Pitch AnythingOren Klaff
Sources 110 entries+

Working bibliography across the research. Inclusion does not mean a source is cited in the proposal.

110 sources

  • Bahmani, S., Park, J. J., Paschalidou, D., Yan, X., Wetzstein, G., Guibas, L. J., & Tagliasacchi, A. (2023). CC3D: Layout-conditioned generation of compositional 3D scenes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 7137–7147). https://doi.org/10.1109/ICCV51070.2023.00659
  • Bar-Tal, O., et al. (2024). Lumiere: A space-time diffusion model for video generation. ArXiv.org.https://arxiv.org/abs/2401.12945
  • Barron, J. T., et al. (2021). Mip-NeRF: A multiscale representation for anti-aliasing neural radiance fields. ArXiv.org. https://arxiv.org/abs/2103.13415
  • Batra, H., Tu, H., Chen, H., Lin, Y., Xie, C., & Clark, R. (2025). SpatialThinker: Reinforcing 3D reasoning in multimodal LLMs via spatial rewards [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.07403
  • Behrens, T. E. J., Muller, T. H., Whittington, J. C. R., Mark, S., Baram, A. B., Stachenfeld, K. L., & Kurth-Nelson, Z. (2018). What is a cognitive map? Organizing knowledge for flexible behavior. Neuron, 100(2), 490–509. https://doi.org/10.1016/j.neuron.2018.10.002
  • Bowker, G. C., & Star, S. L. (1999). Sorting things out: Classification and its consequences. MIT Press.
  • Bruce, J., et al. (2024). Genie: Generative interactive environments. ArXiv.org. https://arxiv.org/abs/2402.15391
  • Chattopadhyay, A., Zhang, X., Wipf, D. P., Arora, H., & Vidal, R. (2023). Learning graph variational autoencoders with constraints and structured priors for conditional indoor 3D scene generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (pp. 785–794). https://doi.org/10.1109/WACV56688.2023.00085
  • Chen, B., Xu, Z., Kirmani, S., Ichter, B., Sadigh, D., Guibas, L., & Xia, F. (2024). SpatialVLM: Endowing vision-language models with spatial reasoning capabilities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 14455–14465). https://doi.org/10.1109/CVPR52733.2024.01370
  • Chen, S., Zhu, T., Zhou, R., Zhang, J., Gao, S., Niebles, J. C., Geva, M., He, J., Wu, J., & Li, M. (2025). Why is spatial reasoning hard for VLMs? An attention mechanism perspective on focus areas. In A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, & J. Zhu (Eds.), Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 9910–9932). PMLR. https://proceedings.mlr.press/v267/chen25cr.html
  • Chen, Y., Nguyen, H. T., Voleti, V., Jampani, V., & Jiang, H. (2025). HouseCrafter: Lifting floorplans to 3D scenes with 2D diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 28440–28450).
  • Cheng, A.-C., Yin, H., Fu, Y., Guo, Q., Yang, R., Kautz, J., Wang, X., & Liu, S. (2024). SpatialRGPT: Grounded spatial reasoning in vision-language models. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 135062–135093). Curran Associates, Inc. https://doi.org/10.52202/079017-4293
  • Chiriatti, M., Ganapini, M., Panai, E., Ubiali, M., & Riva, G. (2024). The case for human–AI interaction as system 0 thinking. Nature Human Behaviour, 8(11), 1995–1997. https://doi.org/10.1038/s41562-024-01995-5
  • Christian, B., & Griffiths, T. (2016). Algorithms to live by: The computer science of human decisions. Henry Holt and Company.
  • Clark, A. (1997). Being there: Putting brain, body, and world together again. MIT Press.
  • Daston, L., & Galison, P. (2007). Objectivity. Zone Books.
  • Deng, A., Cao, T., Chen, Z., & Hooi, B. (2025). Words or vision: Do vision-language models have blind faith in text? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2503.02199
  • Doshi-Velez, F., & Kim, B. (2017). Towards a rigorous science of interpretable machine learning [Preprint]. arXiv. https://arxiv.org/abs/1702.08608
  • Dourish, P. (2001). Where the action is: The foundations of embodied interaction. MIT Press.
  • Durante, Z., et al. (2025). Agent AI: Surveying the horizons of multimodal interaction. ArXiv.org. https://arxiv.org/abs/2504.01512
  • Erdem, U., et al. (2019). Applications of spatial and temporal reasoning in cognitive robotics. Cognitive Robotics Lab.
  • Eslami, S. M. A., et al. (2018). Neural scene representation and rendering. Science, 360(6394), 1204–1210. https://doi.org/10.1126/science.aar6170
  • Feng, M., Hou, H., Zhang, L., Wu, Z., Guo, Y., & Mian, A. (2023). 3D spatial multimodal knowledge accumulation for scene graph prediction in point cloud. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9182–9191). https://doi.org/10.1109/CVPR52729.2023.00886
  • Gao, J., et al. (2021). Dynamic view synthesis from dynamic monocular video. ArXiv.org. https://arxiv.org/abs/2105.06468
  • Garcin, S., Walker, T., McDonagh, S., Pearce, T., Bilen, H., He, T., Wang, K., & Bian, J. (2026). Beyond pixel histories: World models with persistent 3D state [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.03482
  • Gibson, J. J. (1979). The ecological approach to visual perception. Houghton Mifflin.
  • Gibson, J. J. (1979). The theory of affordances. In The ecological approach to visual perception (pp. 127–143).
  • Goodwin, C. (1994). Professional vision. American Anthropologist, 96(3), 606–633.
  • Green, T. R. G., & Petre, M. (1996). Usability analysis of visual programming environments: A ‘cognitive dimensions’ framework. Journal of Visual Languages & Computing, 7(2), 131–174. https://doi.org/10.1006/jvlc.1996.0009
  • Ha, D., & Schmidhuber, J. (2018). World models. ArXiv.org. https://arxiv.org/abs/1803.10122
  • Hafner, D., et al. (2023). Mastering diverse domains through world models. ArXiv.org. https://arxiv.org/abs/2311.01460
  • Hassabis, D., Kumaran, D., Summerfield, C., & Botvinick, M. (2017). Neuroscience-inspired artificial intelligence. Neuron, 95(2), 245–258. https://doi.org/10.1016/j.neuron.2017.06.011
  • Herweijer, C. (2026, July 15). World models are AI's next frontier. TIME.
  • Hu, S., et al. (2024). Simulating the real world: A unified survey of multimodal generative models. ArXiv.org. https://arxiv.org/abs/2405.14034
  • Hu, Y., et al. (2025). EWMBENCH: Evaluating scene, motion, and semantic quality in embodied world models. ArXiv.org. https://arxiv.org/abs/2505.09694
  • Huang, Z., et al. (2024). EnerVerse: Envisioning embodied future space for robotic manipulation. ArXiv.org. https://arxiv.org/abs/2407.08768
  • Hutchins, E. (1995). Cognition in the wild. MIT Press.
  • Jewitt, C. (Ed.). (2009). The Routledge handbook of multimodal analysis. Routledge.
  • Jiang, L., et al. (2024). EnerVerse-AC: Envisioning embodied environments with action-conditioned world models. ArXiv.org. https://arxiv.org/abs/2407.08769
  • Kerbl, B., et al. (2023). 3D Gaussian splatting for real-time radiance field rendering. ArXiv.org. https://arxiv.org/abs/2303.13495
  • Khan, H., & Asif, S. (2026). GenCtrl: A formal controllability toolkit for generative models [Preprint]. arXiv. https://arxiv.org/abs/2601.05637
  • Kirsh, D. (1995). The intelligent use of space. Artificial Intelligence, 73(1–2), 31–68.
  • Kress, G., & van Leeuwen, T. (2001). Multimodal discourse: The modes and media of contemporary communication. Arnold.
  • Kumar, A., et al. (2019). Consistent generative query networks. ArXiv.org. https://arxiv.org/abs/1807.02033
  • Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3706598.3713778
  • Li, C., Peh, E., & Fernando, B. (2025). HMR3D: Hierarchical multimodal representation for 3D scene understanding with large vision-language model [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.22961
  • Li, C., Wu, W., Zhang, H., Xia, Y., Mao, S., Dong, L., Vulić, I., & Wei, F. (2025). Imagine while reasoning in space: Multimodal visualization-of-thought. In A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, & J. Zhu (Eds.), Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 36340–36364). PMLR. https://proceedings.mlr.press/v267/li25cz.html
  • Li, C., Zhang, C., Zhou, H., Collier, N., Korhonen, A., & Vulić, I. (2024). TopViewRS: Vision-language models as top-view spatial reasoners. In Y. Al-Onaizan, M. Bansal, & Y.-N. Chen (Eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 1786–1807). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.106
  • Li, M., Xie, C., Wu, Y., Zhang, L., & Wang, M. (2025). FiVE-Bench: A fine-grained video editing benchmark for evaluating emerging diffusion and rectified flow models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 16672–16681).
  • Liapis, A., Yannakakis, G. N., Alexopoulos, C., & Lopes, P. (2016). Can computers foster human users' creativity? Theory and praxis of mixed-initiative co-creativity. Digital Culture & Education, 8(2), 136–153.
  • Lin, J., Zhu, C., Xu, R., Mao, X., Liu, X., Wang, T., & Pang, J. (2025). OST-Bench: Evaluating the capabilities of MLLMs in online spatio-temporal scene understanding [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.07984
  • Lin, Z., Ehsan, U., Agarwal, R., Dani, S., Vashishth, V., & Riedl, M. (2023). Beyond prompts: Exploring the design space of mixed-initiative co-creativity systems [Preprint]. arXiv. https://arxiv.org/abs/2305.07465
  • Ling, L., Ge, Y., Sheng, Y., & Bera, A. (2025). I-Scene: 3D instance models are implicit generalizable spatial learners [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.13683
  • Linghu, X., Huang, J., Niu, X., Ma, X., Jia, B., & Huang, S. (2024). Multi-modal situated reasoning in 3D scenes. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 140903–140936). Curran Associates, Inc. https://doi.org/10.52202/079017-4473
  • Linghu, X., Huang, J., Zhu, Z., Jia, B., & Huang, S. (2026). SceneCOT: Eliciting grounded chain-of-thought reasoning in 3D scenes. In Proceedings of the International Conference on Learning Representations. https://openreview.net/forum?id=U9meoc0Sau
  • Lipton, Z. C. (2018). The mythos of model interpretability. Queue, 16(3), 31–57. https://doi.org/10.1145/3236386.3241340
  • Liu, R., Wu, R., Van Hoorick, B., Tokmakov, P., Zakharov, S., & Vondrick, C. (2023). Zero-1-to-3: Zero-shot one image to 3D object. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 9264–9275). https://doi.org/10.1109/ICCV51070.2023.00853
  • Liu, Y., et al. (2024). Aligning cyberspace with the physical world: A comprehensive survey on embodied AI. IEEE TPAMI.
  • Liu, Y., Li, X., Zhang, Y., Qi, L., Li, X., Wang, W., Li, C., Li, X., & Yang, M.-H. (2025). Controllable 3D outdoor scene generation via scene graphs. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 28052–28062).
  • Long, X., Guo, Y.-C., Lin, C., Liu, Y., Dou, Z., Liu, L., Ma, Y., Zhang, S.-H., Habermann, M., Theobalt, C., & Wang, W. (2024). Wonder3D: Single image to 3D using cross-domain diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9970–9980). https://doi.org/10.1109/CVPR52733.2024.00951
  • Lu, Y., Luo, W., Tu, P., Li, H., Zhu, H., Yu, Z., Wang, X., Chen, X., Peng, X., Li, X., & Chen, Z. (2025). 4DWorldBench: A comprehensive evaluation framework for 3D/4D world generation models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.19836
  • Lyu, R., Lin, J., Wang, T., Yang, S., Mao, X., Chen, Y., Xu, R., Huang, H., Zhu, C., Lin, D., & Pang, J. (2024). MMScan: A multi-modal 3D scene dataset with hierarchical grounded language annotations. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, & C. Zhang (Eds.), Advances in Neural Information Processing Systems (Vol. 37, pp. 50898–50924). Curran Associates, Inc. https://doi.org/10.52202/079017-1611
  • Ma, W., Chen, H., Zhang, G., Chou, Y.-C., Chen, J., de Melo, C., & Yuille, A. (2025). 3DSRBench: A comprehensive 3D spatial reasoning benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 6924–6934).
  • Martin-Brualla, R., et al. (2021). NeRF in the Wild: Neural radiance fields for unconstrained photo collections. ArXiv.org. https://arxiv.org/abs/2008.02268
  • Mildenhall, B., et al. (2020). NeRF: Representing scenes as neural radiance fields for view synthesis. ArXiv.org. https://arxiv.org/abs/2003.08934
  • Morris, B. (2026, July 2). AI is 'not smart' so what's next in artificial intelligence? BBC News. https://www.bbc.com/news/articles/cj6gr0xkyr3o
  • Naik, S., Shukla, P., Obi, I., Backus, J., Rasche, N., & Parsons, P. (2025). Tracing the invisible: Understanding students' judgment in AI-supported design work. In Creativity and Cognition (C&C '25). ACM. https://doi.org/10.1145/3698061.3734399
  • Nguyen, T., Michaels, J., Fiterau, M., & Jensen, D. (2025). Challenges in understanding modality conflict in vision-language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.02805
  • Nooralahzadeh, F., Rohanian, O., Zhang, Y., Fürst, J., & Stockinger, K. (2026). Arbitration failure, not perceptual blindness: How vision-language models resolve visual-linguistic conflicts [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2604.09364
  • O’Keefe, J., et al. (1978). The hippocampus as a cognitive map. Oxford University Press.
  • Park, J., Jang, K. J., & Tariq, B. (2025). Assessing modality bias in video question answering benchmarks with multimodal large language models. In Proceedings of the AAAI Conference on Artificial Intelligence. https://arxiv.org/abs/2408.12763
  • Park, K., et al. (2021). HyperNeRF: A higher-dimensional representation for topologically varying neural radiance fields. ArXiv.org. https://arxiv.org/abs/2103.16376
  • Poole, B., et al. (2022). DreamFusion: Text-to-3D using 2D diffusion. ArXiv.org. https://arxiv.org/abs/2209.14988
  • Qin, C., et al. (2024). WorldSimBench: Towards video generation models as world simulators. ArXiv.org. https://arxiv.org/abs/2403.12031
  • Rezwana, J., & Maher, M. L. (2023). Designing creative AI partners with COFI: A framework for modeling interaction in human-AI co-creative systems. ACM Transactions on Computer-Human Interaction, 30(5), 1–28. https://doi.org/10.1145/3519026
  • Rombach, R., et al. (2022). High-resolution image synthesis with latent diffusion models. ArXiv.org. https://arxiv.org/abs/2112.10752
  • Sajjadi, M. S. M., et al. (2022). OSRT: Object scene representation transformer. NeurIPS. https://osrt-paper.github.io
  • Sajjadi, M. S. M., et al. (2022). Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. CVPR. https://srt-paper.github.io
  • Sargent, K., Li, Z., Shah, T., Herrmann, C., Yu, H.-X., Zhang, Y., Chan, E. R., Lagun, D., Li, F.-F., Sun, D., & Wu, J. (2024). ZeroNVS: Zero-shot 360-degree view synthesis from a single image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9420–9429). https://doi.org/10.1109/CVPR52733.2024.00900
  • Sarkar, A. (2024). Intension is the new attention: A cognitive perspective on natural language interaction with generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM.
  • Sheng, Z., Han, X., Zhang, Z., Xiong, Z., Ding, Y., Ping, A., Li, X., Guo, T., & Mao, Y. (2026). InEdit-Bench: Benchmarking intermediate logical pathways for intelligent image editing models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.03657
  • Shi, Y., Wang, P., Ye, J., Mai, L., Li, K., & Yang, X. (2024). MVDream: Multi-view diffusion for 3D generation. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=FUgrjq2pbB
  • Sim, M. Y., Zhang, W. E., Dai, X., & Fang, B. (2025). Can VLMs actually see and read? A survey on modality collapse in vision-language models. In Findings of the Association for Computational Linguistics: ACL 2025 (pp. 24452–24470). Association for Computational Linguistics. https://aclanthology.org/2025.findings-acl.1256/
  • Spelke, E. S., et al. (2007). Core knowledge. Developmental Science. https://doi.org/10.1111/j.1467-7687.2007.00569.x
  • Stogiannidis, I., McDonagh, S., & Tsaftaris, S. A. (2025). Mind the gap: Benchmarking spatial reasoning in vision-language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2503.19707
  • Suchman, L. A. (1987). Plans and situated actions: The problem of human-machine communication. Cambridge University Press.
  • Sun, Y., Jang, E., Ma, F., & Wang, T. (2024). Generative AI in the wild: Prospects, challenges, and strategies. In Proceedings of the CHI Conference on Human Factors in Computing Systems (pp. 1–16). ACM. https://arxiv.org/abs/2404.04101
  • Suresh, H., & Guttag, J. (2021). A framework for understanding sources of harm throughout the machine learning life cycle. In Equity and Access in Algorithms, Mechanisms, and Optimization (EAAMO '21) (pp. 1–9). ACM. https://doi.org/10.1145/3465416.3483305
  • Tankelevitch, L., Kewenig, V., Simkute, A., Scott, A. E., Sarkar, A., Sellen, A., & Rintel, S. (2024). The metacognitive demands and opportunities of generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems. ACM. https://doi.org/10.1145/3613904.3642902 · candidate 260814
  • Tewari, A., et al. (2020). State of the art on neural rendering. ArXiv.org. https://arxiv.org/abs/2004.04776
  • Tversky, B. (2011). Visualizing thought. Topics in Cognitive Science, 3(3), 499–535. https://doi.org/10.1111/j.1756-8765.2010.01113.x
  • Ullman, S. (1984). Visual routines. Cognition.
  • Van de Maele, T., Dhoedt, B., Verbelen, T., & Pezzulo, G. (2024). A hierarchical active inference model of spatial alternation tasks and the hippocampal-prefrontal circuit. Nature Communications, 15, Article 9892. https://doi.org/10.1038/s41467-024-54257-3
  • Wang, C., Zhou, Y., Wang, Q., Wang, Z., & Zhang, K. (2025). ComplexBench-Edit: Benchmarking complex instruction-driven image editing via compositional dependencies. In Proceedings of the 33rd ACM International Conference on Multimedia (pp. 13391–13397).
  • Wang, Q., Zhang, Y., Holynski, A., Efros, A. A., & Kanazawa, A. (2025). Continuous 3D perception model with persistent state. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. https://arxiv.org/abs/2501.12387
  • Wang, Z., et al. (2025). SITE: Towards spatial intelligence through evaluation. ArXiv.org. https://arxiv.org/abs/2503.09733
  • Whittington, J. C. R., Muller, T. H., Mark, S., Chen, G., Barry, C., Burgess, N., & Behrens, T. E. J. (2020). The Tolman-Eichenbaum machine: Unifying space and relational memory through generalization in the hippocampal formation. Cell, 183(5), 1249–1263.e23. https://doi.org/10.1016/j.cell.2020.10.024
  • Wu, G., et al. (2024). 4D Gaussian splatting for real-time dynamic scene rendering. CVPR. https://guanjunwu.github.io/4dgs
  • Wu, M., Ji, J., Huang, O., Li, J., Wu, Y., Sun, X., & Ji, R. (2024). Evaluating and analyzing relationship hallucinations in large vision-language models. In R. Salakhutdinov, Z. Kolter, K. Heller, A. Weller, N. Oliver, J. Scarlett, & F. Berkenkamp (Eds.), Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 53553–53570). PMLR. https://proceedings.mlr.press/v235/wu24l.html
  • Wu, X., Liang, D., Feng, T., Xia, K., Zhang, Y., Li, X., Tan, X., & Bai, X. (2026). Generation models know space: Unleashing implicit 3D priors for scene understanding [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2603.19235
  • Yang, J., Yang, S., Gupta, A. W., Han, R., Li, F.-F., & Xie, S. (2025). Thinking in space: How multimodal large language models see, remember, and recall spaces. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 10632–10643). https://doi.org/10.1109/CVPR52734.2025.00994
  • Yang, M., et al. (2024). Cambrian-S: Towards spatial supersensing in video. ArXiv.org. https://arxiv.org/abs/2403.12843
  • Yannakakis, G. N., Liapis, A., & Alexopoulos, C. (2014). Mixed-initiative co-creativity. In Proceedings of the 9th International Conference on the Foundations of Digital Games.
  • Yu, A., et al. (2021). PlenOctrees for real-time rendering of neural radiance fields. ArXiv.org. https://arxiv.org/abs/2103.14024
  • Yuan, D., Graham, C., Yakura, H., & Lim, A. (2026). Using the blackbox in embodied AI art practice: Uncertainty at the interface. In Designing Interactive Systems Conference (DIS ’26) (pp. 4181–4201). ACM. https://doi.org/10.1145/3800645.3813020
  • Zhang, W., Huang, Y., Xu, Y., Huang, J., Zhi, H., Ren, S., Xu, W., & Zhang, J. (2025). Why do MLLMs struggle with spatial understanding? A systematic analysis from data to architecture [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2509.02359
  • Zhang, Z., Zhou, W., Zhao, J., & Li, H. (2025). Robust multimodal large language models against modality conflict. In Proceedings of the 42nd International Conference on Machine Learning. https://arxiv.org/abs/2507.07151
  • Zhao, R., Zhang, Z., Xu, J., Chang, J., Chen, D., Li, L., Sun, W., & Wei, Z. (2025). SpaceMind: Camera-guided modality fusion for spatial reasoning in vision-language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.23075
  • Zhao, X., et al. (2024). Pseudo-generalized dynamic view synthesis from a video. ICLR. https://arxiv.org/abs/2310.08587
  • Zhen, X., et al. (2024). 3D-VLA: A 3D vision-language-action generative world model. ArXiv.org. https://arxiv.org/abs/2403.15952

August 2026 · Archived prototype

August 29 presentation figure: inputs supplied versus recurrence of inference, with highlighted Photoshop operations

29 August 2026 · Presentation figure