without specialised sensors or app installation. The methodology combines pretrained models rather than
training a new network: straight-line and monocular depth maps are extracted from the photograph as
structural features and used to condition a ControlNet-guided diffusion model; the redesigned image is
converted into a full-room mesh by pinhole back-projection of a monocular depth estimate; and the mesh is
delivered to the browser's AR runtime. The specific contributions are:
(1) A geometry-preserving redesign pipeline in which M-LSD line maps
[19]
and Depth-Anything-V2 depth
maps
[20]
serve as structural conditions for ControlNet,
[16]
embedded in a four-layer provider cascade that
maintains availability on free-tier inference infrastructure (Algorithm 1).
(2) A single-image, full-room reconstruction procedure that converts a depth map into a colour-textured,
artefact-filtered GLB mesh (Algorithm 2), with closed-form expressions for the mesh size and storage
footprint (Eqs. (3) – (9)).
(3) An installation-free delivery path from photograph to AR through WebXR, Android Scene Viewer and iOS
Quick Look, backed by row-level data isolation between users.
(4) An evaluation on real user requests comprising structure-preservation, style-adherence (with a per-class
classification report), latency, reliability and mesh metrics, including a comparison against unconditioned
text-only generation.
The contribution is mainly in system design and integration. Every model in the pipeline is an existing
pretrained component; what is new is how the models are combined, constrained and delivered for the specific
job of previewing a redesign of an existing room. Section 2 reviews related work; Section 3 covers the system
overview, working principle, software, implementation, model configuration and testing protocol; Section 4
reports the results, a competitive analysis, a comparison of methods, an interpretation and the limitations; and
Section 5 concludes.
2. Literature review
2.1 Augmented reality in interior design
Revathy et al. reviewed the use of AR in interior design and its capacity to let users try furniture arrangements
in real environments.
[6]
Samant and Vartak
[2]
and Juned et al.
[7]
developed mobile AR applications in which users
place virtual furniture models in their room, the latter using ARCore with a cloud-hosted model catalogue, and
Huang and Ni presented ARID, an AR interior design system.
[17]
Dave et al. examined how AR, virtual reality (VR)
and AI influence consumer decisions in housing design.
[1]
Outside the home-design domain, Zimmermann et al.
showed in an online experiment that an AR shopping assistant with personalised, explainable recommendations
improved perceived in-store shopping experience,
[8]
and Kim et al. studied how uncertainty in AI-based object
detection should be exposed to users in AR interaction design.
[21]
These systems demonstrate the value of in-
situ visualisation, but they generally operate on individual catalogue assets rather than on a restyled rendering
of the entire existing room.
2.2 AI for interior design and layout
Cuevas et al. presented LayOut Loud, an AI-powered AR and mobile application for room interior design and
layout optimisation.
[9]
K n et al. proposed an algorithm based on a hierarchical tree of procedural rules that
generates personalised furniture configurations and presents them in mobile AR and evaluated it in three user
studies.
[10]
Wu and Han evaluated the combination of AI and VR for interactive interior decoration;
[11]
Gong
[5]
applied AI-based visual feature modelling to interior design schemes; Jiang
[12]
built a generative adversarial
network for personalised furniture design within an Internet of Things environment; Samuel et al.
[13]
discussed
computational technology for design graphics and modelling; and Almusaed et al. reviewed AI models for smart-
home living spaces.
[14]
Broader reviews cover AI, VR and AR in product design
[3]
and the convergence of AI and
AR.
[4]
The common output of these approaches is an arrangement, selection or design of discrete furniture
objects; photorealistic restyling of the user's own photograph under explicit geometric constraints is not their
focus.
2.3 Controllable image generation
Latent diffusion models perform the denoising process in the latent space of an autoencoder, which makes high-
resolution text-to-image synthesis tractable;
[15]
SDXL scales this approach to higher resolutions.
[18]
InstructPix2Pix edits an input image according to a natural-language instruction
[22]
but does not explicitly
enforce structural constraints. ControlNet attaches a trainable copy of the diffusion encoder to a frozen
pretrained model and injects features computed from a control image, such as edges, line segments or depth,
into the denoising network.
[16]
M-LSD is a lightweight line-segment detector designed for real-time use on
constrained devices;
[19]
because interior architecture is dominated by straight lines, its output is a natural
control signal for preserving walls and openings
2.4 Monocular depth estimation and single-image 3D