Quantumania 11 - Daggers, Bras and Kets
By this point you could be forgiven for wondering what any of this has to do with QM. Let’s bring in complex numbers. If our vector space has complex scalars, then our matrices that operate on it can have complex entries. In terms of direct visualisation of the impact of such a matrix, good luck with that, but broadly speaking, it should be clear that we can still write a diagonal matrix and that it will scale along the coordinate axes. Likewise we can apply rotations:
\[U = \begin{bmatrix} a & -b^* \\ b & a^* \end{bmatrix} \quad \text{where } \vert a \vert^2 + \vert b \vert^2 = 1\]Note that this time I haven’t jumped straight to writing it with entries $\cos \theta$ etc. I’ve just given the constraint about the squared-moduli of $a$ and $b$ summing to $1$.
This has to do with one of the defining characteristics of a rotation, which is that the vectors all stay the same length.
Pythagoras computes the length of a vector from its coordinates. It looks like $a$ and $b$ are acting like vector coordinates in that constraint, except they’re complex, so we’re taking the modulus squared.
There’s a way to get the modulus squared of a complex number, $z$, which is to multiply it by its complex conjugate, $z^*$, which we get by flipping the sign of the imaginary part:
\[z = p + qi\] \[z^* = p - qi\] \[z z^* = (p + qi)(p - qi) = p^2 - pqi + pqi - q^2i^2 = p^2 + q^2\]Look how they Pythagorised my boy! There is nothing surprising about this: flipping the sign of $qi$ is like reflecting in the real axis, so if you think of a complex number as a rotation factor, its conjugate rotates the other way. So conjugation on the polar form is like this:
\[z = r e^{i\theta}\] \[z^* = \left(r e^{i\theta}\right)^* = r e^{-i\theta}\]So multiplying a complex number by its own conjugate rotates it back to the real number line, cancelling out the rotation, and guaranteeing the result is a real number, the square of the modulus:
\[z z^* = r e^{i\theta} \left(r e^{i\theta}\right)^* = r^2 e^{i\theta} e^{-i\theta} = r^2\]So this is how we get a real, non-negative “length” of an individual complex number. But in our complex vector spaces, we’d like the length of a vector to also be real and non-negative. Let’s write Pythagoras for the whole vector $\vec{v}$ with two complex coordinates $(v_1, v_2)$:
\[\Vert \vec{v} \Vert^2 = v_1 v_1^* + v_2 v_2^*\]Now, there’s a massively important thing about our vector spaces we’ve got this far without mentioning, but here it comes: the inner product. This is a general term for the same thing as the dot product we use with regular vectors.
This is a function that takes two vectors and outputs a scalar, and it is the essential link between a dry column-vector space like $R^2$ and the world of geometry. The inner product of two orthogonal vectors is $0$. And crucially the inner product of a vector with itself is the squared length of the vector. Now, looking at the above formula for the length of the vector, it strongly suggests that the inner product of two vectors $\vec{v}$ and $\vec{w}$ should be:
\[(\vec{v}, \vec{w}) = v_1 w_1^* + v_2 w_2^*\]simply because if we substitute $\vec{v}$ in place of $\vec{w}$ then we’ll get back our length-squared formula, as we must. This is the same thing as taking the complex conjugate of all the coordinates of one vector, before applying the usual dot product formula we’re familiar with from real vector spaces.
Also we can rewrite the dot product as matrix multiplication if we transpose the left vector from a column to a row. So that we put all this fiddling around onto one of the vectors instead of sharing it out, let’s define the superscripted symbol $\dagger$ to mean:
- transpose the matrix, and
- take the complex conjugate of all its entries.
So we can just define the inner product of the column vectors $v$ and $w$ as $v^\dagger$ $w$. The left column vector $v$ is transposed into a row, all its entries are conjugated, and then ordinary matrix multiplication with $w$ does the rest.
I’ve mostly been writing vectors in the strange QM way: $\vert \psi \rangle$. This is really one half of the notation invented by Dirac. The other half is:
\[\langle \psi \vert \equiv \psi^\dagger\]In other words, $\vert \psi \rangle$ is an abstract vector plucked from the space. Then $\psi$ is a concrete representation of it in a particular basis as a column of numbers. By applying the $^\dagger$ operator to it, we turn it into a row containing the complex-conjugated coordinates of $\psi$. And we can take this as being the concrete representation of a fundamentally different object, $\langle \psi \vert$, albeit one that is paired uniquely with $\vert \psi \rangle$.
What is the nature of this new object? It is in fact (unsurprisingly, given its concrete representation) another kind of vector, from a different vector space, but it can be thought of a function that acts on the original, ordinary type of vector. The terminology in this area is so varied you can take your pick of names, but in physics they call $\langle \psi \vert$ a bra and $\vert \psi \rangle$ a ket. When a bra (function) acts on a ket (input), we have a bra-ket (Dirac thought this was funny):
\[\langle \psi \vert \psi \rangle\]In that example, as both sides say $\psi$, and assuming that’s a state vector, the answer should be 1, because the result is the squared length, and it better be a unit vector.
It has a double meaning: in abstract, it’s a sort of geometrical function acting on a vector, a geometrical object, to produce a scalar. In concrete terms, when we’ve expressed both sides in compatible bases as row and column vectors, it’s matrix multiplication.
So when we talk about the inner product between two vectors in QM, what we really mean is that we have two kets, and we’re going to convert one of them into a bra in the standard way, and let that bra act on the ket. This is most often assumed without stating it explicitly, so you might say you’re going to take the inner product of $\vert a \rangle$ and $\vert b \rangle$ and so that’s $\langle a \vert b \rangle$. It’s a given that you’re going to have switch one of them to a bra.
Does it matter which? Yes: these are mutual complex conjugates:
\[\langle b \vert a \rangle = \langle a \vert b \rangle^*\]Whenever you see that bra-ket shape, you’re looking at a scalar. The most helpful way to think about the purpose of a bra is to think of a basis vector, such as $\vert u \rangle$, and suppose that you have a state vector $\vert \psi \rangle$ and you want to know its coordinate in the direction of $\vert u \rangle$. It’s just:
\[\langle u \vert \psi \rangle\]So I can make a column vector of the coordinates in the up/down basis:
\[\begin{bmatrix} \langle u \vert \psi \rangle \\ \langle d \vert \psi \rangle \\ \end{bmatrix}\]And I can (somewhat unnecessarily) reconstruct the whole vector by scaling the basis vectors by the coordinates in that basis:
\[\vert \psi \rangle = \langle u \vert \psi \rangle \vert u \rangle + \langle d \vert \psi \rangle \vert d \rangle\]I’m scaling the two basis vectors by the coordinates, but I’ve got those coordinates by fetching them out of $\vert \psi \rangle$. Of course, if $\vert \psi \rangle$ is already known to me as a column vector in that basis, then fetching a coordinate from it isn’t that arduous. It’s right there, in the column vector! But the neat thing about this geometrically is that we can work in any basis. If I have $\vert \psi \rangle$ and the up/down basis vectors as column vectors all expressed in the front/back basis, the above expressions all produce the same results.
Let’s return to our general rotation matrix and look at how it works. It’s all the way up the page by now so here it is again:
\[U = \begin{bmatrix} a & -b^* \\ b & a^* \end{bmatrix} \quad \text{where } \vert a \vert^2 + \vert b \vert^2 = 1\]You may recall that we can understand what a matrix does to the unit vectors of its basis by simply reading the columns. The first column is the result it produces when it acts on $(1, 0)^\intercal$, the second column is the same for $(0, 1)^\intercal$.
I mentioned earlier that one of the defining characteristics of a rotation is that the vectors all stay the same length. What I didn’t mention is that it must also preserve the angle between pairs of vectors. Now, by definition our unit vectors are orthogonal before we rotate them:
\[\langle u \vert d \rangle = 0\]After we rotate them they should be the same length and still be orthogonal.
Reading the columns of the matrix, after we rotate them the vectors have coordinates $(a, b)$ and $(-b^\ast, a^\ast)$. So the inner product is:
\[-ab + ba = 0\]They’re still orthogonal, mission accomplished. But are they both still unit vectors? Easy enough to find out, because we just take the inner product of each with itself:
\[a^2 + b^2 \quad \text{and} \quad b b^* + a a^*\]The only way to ensure those both come out as 1 is to require that:
\[\vert a \vert^2 + \vert b \vert^2 = 1\]I mentioned yesterday that for a rotation (or reflection) matrix in a real vector space we can get its inverse by simply transposing it. In a complex vector space we have to upgrade this to using the full $^{\dagger}$ approach:
\[U^\dagger = U^{-1}\]It’s easy to check that this works, because if you multiply a matrix by its inverse, that should give you $I$, the identity matrix, and so:
\[UU^\dagger = I\]So we just need to multiply our general $U$ by $U^\dagger$ and see what we get:
\[UU^{\dagger} = \begin{bmatrix} a & -b^* \\ b & a^* \end{bmatrix} \begin{bmatrix} a^* & b^* \\ -b & a \end{bmatrix} = \begin{bmatrix} aa^* + bb^* & ab^* - ab^* \\ ba^* - ba^* & bb^* + aa^* \\ \end{bmatrix} = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ \end{bmatrix}\]In the last step, we’re using the fact that our general rotation matrix stipulates that $\vert a \vert^2 + \vert b \vert^2 = 1$, and the previously noted rule that $aa^* = \vert a \vert$.
As well as rotating Stick Man, we also applied a kind of orthogonal scaling operation to him, which:
- is represented by a symmetric matrix (equal to its own transpose),
- picks out a set of orthogonal directions in the vector space along which it scales, suitable for basis vectors to point along (its eigenbasis) and
- can always be represented by a diagonal matrix, simply by writing its own eigenbasis.
There’s one aspect of this we need to upgrade for complex vector spaces. We have a reason in QM for wanting the scaling factors to be restricted to real numbers, so when stated in its simple diagonal form, the entries $x$, $y$ on the diagonal are real numbers:
\[H = \begin{bmatrix} x & 0 \\ 0 & y \\ \end{bmatrix}\]As we noted when stretching/squishing Stick Man, by putting the diagonal form inside a rotation sandwich, we get a matrix that scales along any set of orthogonal directions we like:
\[\begin{bmatrix} a & -b^* \\ b & a^* \end{bmatrix} \begin{bmatrix} x & 0 \\ 0 & y \\ \end{bmatrix} \begin{bmatrix} a^* & b^* \\ -b & a \end{bmatrix} = \begin{bmatrix} a & -b^* \\ b & a^* \end{bmatrix} \begin{bmatrix} xa^* & xb^* \\ -yb & ya \\ \end{bmatrix} = \begin{bmatrix} xaa^* + ybb^* & xab^* - yab^* \\ xa^*b - ya^*b & xbb^* + yaa^* \\ \end{bmatrix}\]Now because $x$ and $y$ are real, and $aa^*$ is the modulus of $a$ and therefore real, the main diagonal entries are real numbers.
Also because real numbers have no imaginary part to negate, they are their own complex conjugates, which means that the off-diagonal entries are mutually complex conjugates:
\[xab^* - yab^* = (xa^*b - ya^*b)^*\]What that means is that the general form of an orthogonal scaling matrix is not that it is equal to its own transpose, but that it is equal to its conjugate transpose:
\[H = H^{\dagger}\]That’s the only difference from before. (In fact this is true of the real case too, because the conjugate transpose is equal to the transpose if all the entries in the matrix are real.)
Some updated terminology:
- The conjugate transpose, indicated by $^{\dagger}$, is often called the hermitian conjugate.
- A matrix that is equal to its hermitian conjugate is called a hermitian matrix, and represents a hermitian operator (which is why I used $H$ for it in the example above.)
- A rotation or reflection matrix such as $U$ is called a unitary matrix (which is why I labelled it $U$ above), representing a unitary operator, which preserves the inner product between any pairs of vectors it is applied to.
So now we’ve identified a couple of types of operator and their matrix representations, and it’s high time we found out what they have to do with QM.
Not yet regretting the time you've spent here?
Keep reading: