Combine Multiple Projectors for 3D texturing

Hello,
I 've been using Unity quite sometime but I am just starting with shaders. I am nearing the end of a project and I am stuck at the final part:

-I am getting video from 2 cameras, sending them to unity and by calibrating them, use properly translated/rotated objects with projectors on them and their texture is updated constantly with the camera feed.
-I have an uncolored 3d Mesh, which is also updated and a “faithful” representation of the scene.

What I was hoping to do, was by correctly positioning 2 projectors, to properly map the images on the mesh.
However, using 2 projector/multiply shaders doesn’t give the expected results and I think that 2 projectors can’t do what I want.

The question is : How would I , provided I have 2 textures and 2 point of views (probably using one shader) find the correct color on the mesh? (alpha adds up to 1 always, color is a combination of the camera textures, weighted by their incident angle - if a camera is looking directly at the mesh / parallel to normal, it should obviously have a bigger weight than a camera looking almost parallel to the surface-).

I know this is a super specific question, but I am really out of options here.

bump?

Look at how a projector shader works. It gets a projection matrix passed to it from the projector and uses that to determine the UVs for the projected texture. You could do this in a single shader by manually passing the two projection matrices of the camera views (more specifically the view & projection matrices) to determine the UVs. You can also pass the world position of the two cameras to the shader to easily determine the incident angle. Calculate the world position and world normal of the mesh vertices and then dot(normalize(worldNormal), normalize(worldPosition - cameraPosition)) in the fragment shader.

However this might still not give you what you want since occlusion will mean just using the best incident angle won’t necessarily give you the best data.

Thanks for the reply. I looked into the Projection shader during the weekend and I was thinking along the lines of what you suggested. If I get to the incident angle it should be easy(easier?) from there.

But The projector shader uses the the reverse idea : Finds the vertex it hits, transforms the position to screen space,grabs the uv and the texture color and does Blend DstColor .

It is using a lot of unity helper functions, some which make sense (UnityObjectToClipPos) but stuff like
unity_Projector are not explained. Also these functions take 1 argument, because it is assumed we always do the calculation based on the projector this shader is attached to.

I know how to :
-update the textures
-pass the camera 4x4 or rotation+position vectors for a couple of cameras(that only rarely get changed onEvent).
-How I would combine the color values given angle and color.

BIG question is : How do I calculate the uv coords for a camera-like object?
I assume the complex shader is on the material of the mesh, but what do I put on the objects? Camera? Projector? Just a script with a Texture variable? The projector component takes care of “sending” the fragment color for its attached texture but how do I “grab it”?

Sorry for the wall of text, if it’s too complex/long to answer, even pointing me to a related tutorial/similar shader will do.
I had a look today at Triplanar Mapping, but I don’t know if it’s adaptable to my scenario.

Quick correction. There are no “hit” vertices. The projector is doing an intersection test of the projector’s frustum and the bounds of meshes in the scene and rendering any mesh it overlaps with an additional time with the projector’s material. The vertices are transformed into the camera’s clip space and into the projector’s screen space. Clip space (aka projection space), is a little different than screen space, and you’ll have to understand that a bit. Applying a camera’s projection matrix to an position that is in view space will put it into clip space.

That UnityObjectToClipPos function applies the full MVP (Model View Projection) matrix, which transforms a point from the local object space mesh position to the clip position, hence the name. This diagram may help:


source: https://twitter.com/capnramses/status/936393671519494151

Note the names of each “space” isn’t rigidly defined, so they’re called different things by different people or in different places. Even within Unity’s shader code there’s not a ton of consistency. However it’s good to note that the names of the matrices (on the arrow lines) in the above diagram match up well with Unity’s usage of “Model”, “View”, and “Projection”. An MVP matrix is the model, view, and projection matrix multiplied together so that transforming from local or object space into clip space can be done in a single matrix and vector multiplication.

edit: If you’re curious about the “normalized device space” and “viewport space” in the above diagram and why that doesn’t seem to happen in the shader, it’s because the GPU does that on it’s own. The GPU expects a clip space position output from the vertex shader and does the transformations into viewport space as part of rasterization.

The unity_Projector is essentially a MVP matrix for the projector, but setup so a position is transformed into a “UV” space (essentially 0.0 to 1.0 range “screen space” for the projector).

The unity_ProjectorClip matrix appears to be a similar matrix, but is always an orthographic projection. The reason for this is the depth in clip / projection space is not linear (0.5 is not half way between the near and far clip planes in world space), but for orthographic projections it is linear in world space. You probably won’t need this.

So what you need is the “MVP” for each of your camera views so you can get it into clip space. However it’ll be a lot easier if you just pass the VP matrix so you don’t have to worry about the mesh’s local to world matrix. You can then transform that into “UV” space fairly easily in the shader using the same way Unity does to get screen space positions rather than trying to figure out how to get the matrix setup properly.

psuedo shader code:
// vertex shader
float4 worldPos = mul(unity_ObjectToWorld, v.vertex);
float4 camA_ClipPos = mul(_CameraA_MATRIX_VP, worldPos);
float4 camB_ClipPos = mul(_CameraB_MATRIX_VP, worldPos);
o.camAPos = float4(camA_ClipPos.xy * 0.5 + camA_ClipPos.w * 0.5, camA_ClipPos.zw); // magic!
o.camBPos = float4(camB_ClipPos.xy * 0.5 + camB_ClipPos.w * 0.5, camB_ClipPos.zw);

// fragment shader
float2 camAUV = i.camAPos.xy / i.camAPos.w;
float2 camBUV = i.camBPos.xy / i.camBPos.w;

In C# you can calculate the MATRIX_VP by doing this (assuming you setup dummy cameras objects to align to the camera views instead of projectors):
material.SetMatrix("_CameraA_MATRIX_VP", camA.projectionMatrix * camA.worldToCameraMatrix);

As for getting the incident angle, it’s possible to do it with the projection matrix, but it requires transforming the vertex normal into projection space, which with the above setup requires transforming into world space first, so why do that extra work?

// vertex shader
o.worldNormal = UnityObjectToWorldNormal(v.normal);
o.worldPos = worldPos;

// fragment shader
float incidentA = dot(normalize(i.worldPos.xyz - _CameraAWorldPos.xyz), normalize(i.worldNormal));
float incidentB = dot(normalize(i.worldPos.xyz - _CameraBWorldPos.xyz), normalize(i.worldNormal));

First of all, thank you VERY much for the detailed response! That diagram explained some concepts that I didn’t know or perceived wrongly. My SO saw the thorough response and said: “These people exist, and they are the reason I still have faith in humanity” :slight_smile:

Also the pseudo shader code appears to be almost all the proper code I need, so by applying it in an unlit shader template This is what it looks like:

Shader "Custom/UnlitBasicModded"
{
    Properties
    {
        _TextureA(" Texture A - Left(RGB)", 2D) = "white" {}
        _TextureB(" Texture B - Right(RGB)", 2D) = "white" {}
    }
    SubShader
    {

        Pass
        {
            CGPROGRAM
// Upgrade NOTE: excluded shader from DX11; has structs without semantics (struct v2f members camAPos,camBPos)
#pragma exclude_renderers d3d11

            #pragma vertex vert
            #pragma fragment frag

            #include "UnityCG.cginc"

            struct appdata
            {
                float4 vertex : POSITION;
                float3 normal: NORMAL;
                float2 uv : TEXCOORD0;
            };

            struct v2f
            {
                float4 worldPos : SV_POSITION;
                float3 worldNormal: NORMAL;
                float4 camAPos;
                float4 camBPos;
            };


            sampler2D _TextureA;
            sampler2D _TextureB;
            float4x4 _CameraA_MATRIX_VP;
            float4x4 _CameraB_MATRIX_VP;
            float4 _CameraAWorldPos;
            float4 _CameraBWorldPos;

            v2f vert (appdata v)
            {
                float4 worldPos = mul(unity_ObjectToWorld, v.vertex);
                float4 camA_ClipPos = mul(_CameraA_MATRIX_VP, worldPos);
                float4 camB_ClipPos = mul(_CameraB_MATRIX_VP, worldPos);
                v2f o;
                o.worldNormal = UnityObjectToWorldNormal(v.normal);
                o.camAPos = float4(camA_ClipPos.xy * 0.5 + camA_ClipPos.w * 0.5, camA_ClipPos.zw);
                o.camBPos = float4(camB_ClipPos.xy * 0.5 + camB_ClipPos.w * 0.5, camB_ClipPos.zw);
                o.worldPos = worldPos;
                return o;
            }
           
            fixed4 frag (v2f i) : SV_Target
            {
                float2 camAUV = i.camAPos.xy / i.camAPos.w;
                float2 camBUV = i.camBPos.xy / i.camBPos.w;

                float incidentA = max(0,   dot(normalize(i.worldPos.xyz - _CameraAWorldPos.xyz), normalize(i.worldNormal)));
                float incidentB = max(0, dot(normalize(i.worldPos.xyz - _CameraBWorldPos.xyz), normalize(i.worldNormal)));
                // sample the texture
                fixed4 colA = tex2D(_TextureA, camAUV);
                fixed4 colB =tex2D(_TextureB, camBUV);
                fixed4 col = (colA * incidentA +colB * incidentB) / (incidentA+incidentB);
                return col;
            }
            ENDCG
        }
    }
}

and the script that handles the assignment is this:

public class ShaderInfo : MonoBehaviour {

    public GameObject CamA; // objects with cameras, positioned like real cameras
    public GameObject CamB;
    public Texture texA; //textures from actual cameras
    public Texture texB;

    public Material shaderMaterial;
   
    // Use this for initialization
    void Start () {

        shaderMaterial.SetMatrix("_CameraA_MATRIX_VP", CamA.GetComponent<Camera>().projectionMatrix * CamA.GetComponent<Camera>().worldToCameraMatrix);
        shaderMaterial.SetMatrix("_CameraB_MATRIX_VP", CamB.GetComponent<Camera>().projectionMatrix * CamB.GetComponent<Camera>().worldToCameraMatrix);

        shaderMaterial.SetTexture ("_TextureA", texA);
        shaderMaterial.SetTexture ("_TextureB", texB);
        Vector4 positionA = new Vector4 (CamA.transform.position.x, CamA.transform.position.y, CamA.transform.position.z, 1f);
        Vector4 positionB = new Vector4 (CamB.transform.position.x, CamB.transform.position.y, CamB.transform.position.z, 1f);
        shaderMaterial.SetVector("_CameraAWorldPos", positionA);
        shaderMaterial.SetVector("_CameraBWorldPos",positionB);

    }
   
}

Now the obvious problem I can find here (there may be more) is that in order to pass the per-pixel camA and camB positions, they need to be defined in the v2f struct (since it sends vector to fragment data).
Variables in the v2f struct need to have semantics (thats why it keeps adding exclude_renderers d3d11), but from reading the documentation, semantics are vertex related data like position,normal,uv etc.
I tried using :TEXCOORD5 and :TEXCOORD6 which are also float4, but it just goes from bring pink to transparent.

Semantics are used to pass data to the vertex shader (in the form of the “appdata” struct), from the output of the vertex shader to the input of the fragment shader (in the “v2f” struct), and for the output of the fragment shader (usually just SV_Target on the function definition). Most of these can be any arbitrary data, but each stage has a limited number of valid semantics that can be used, and some have special purposes. For example the SV_POSITION output semantic of the vertex shader is one of the few that must be specific data, it must be the clip space position of the vertex. Otherwise anything passed via a TEXCOORD# or COLOR# is just getting passed along as the data it is, interpolated . There’s a limit on the number of those semantics you can use, not for each type, but in total. However that limit is 16 float4 values (or 64 values in total), so TEXCOORD5 and TEXCOORD6 should be fine, though I try to start at zero out of habit.

Technically NORMAL is not a valid vertex shader output semantic, but Unity allows it. Not sure if this is something that’s being fixed on Unity’s side or in the shader compiler.

The TLDR version is just use the TEXCOORD# semantic, starting at 0.

            struct v2f
            {
                float4 worldPos : SV_POSITION;
                float3 worldNormal : TEXCOORD0;
                float4 camAPos : TEXCOORD1;
                float4 camBPos : TEXCOORD2;
            };

However that still won’t work because, as I mentioned, SV_POSITION is a magic semantic that must be the clip space position. You’re currently using that semantic to pass the world position, which is going to be totally wrong even if the rest of the shader is working. Basically, it’s going “transparent” because you’re passing the wrong data to the SV_POSITION. Unless the mesh you’re rendering is less than 2 units wide and centered around 0,0,0 in world space it would not end up as a valid clip space position and be culled.

            struct v2f
            {
                float4 pos : SV_POSITION; // must be clip space pos!
                float4 worldPos : TEXCOORD0;
                float3 worldNormal : TEXCOORD1;
                float4 camAPos : TEXCOORD2;
                float4 camBPos : TEXCOORD3;
            };

            v2f vert (appdata v)
            {
                v2f o;

                float4 worldPos = mul(unity_ObjectToWorld, v.vertex);
                float4 camA_ClipPos = mul(_CameraA_MATRIX_VP, worldPos);
                float4 camB_ClipPos = mul(_CameraB_MATRIX_VP, worldPos);

                o.pos = UnityObjectToClipPos(v.vertex);
                o.worldNormal = UnityObjectToWorldNormal(v.normal);
                o.camAPos = float4(camA_ClipPos.xy * 0.5 + camA_ClipPos.w * 0.5, camA_ClipPos.zw);
                o.camBPos = float4(camB_ClipPos.xy * 0.5 + camB_ClipPos.w * 0.5, camB_ClipPos.zw);
                o.worldPos = worldPos;
                return o;
            }