Notation is compressed language

Every formula is a sentence written in shorthand. Reading it well means expanding the shorthand: saying the formula aloud in words, knowing which symbols are fixed and which vary, and knowing which conventions the author is relying on without stating them. Two questions do most of the work.

  • Which symbols are bound? A bound (dummy) variable, such as xx in ∫01x2 dx\int_0^1 x^2\,dx or kk in ∑k=1nk\sum_{k=1}^{n} k, can be renamed without changing the meaning. The expression does not depend on it.
  • Which symbols are free? The result depends on these. In ∑k=1nk=n(n+1)/2\sum_{k=1}^{n} k = n(n+1)/2, the free symbol is nn.

When a formula confuses you, list its free symbols. Often the confusion is that you were treating a dummy variable as if the answer depended on it, or the reverse.

Derivative operators

The notations dydx\frac{dy}{dx}, y′y', DyDy and, in physics, y˙\dot y (for time derivatives) all name the same derivative. Operator notation D=ddxD = \frac{d}{dx} is useful because operators can be added, multiplied by constants and composed. Powers mean repeated application: D2y=y′′D^2 y = y''. For constant coefficients, operator polynomials multiply like ordinary polynomials:

(D−1)(D−2) y=(D−1)(y′−2y)=y′′−3y′+2y=(D2−3D+2) y(D-1)(D-2)\,y = (D-1)(y' - 2y) = y'' - 3y' + 2y = (D^2 - 3D + 2)\,y

An inverse operator such as (D+b)−1(D+b)^{-1} means "the operation that undoes D+bD + b." It is only meaningful once a space of functions is fixed on which D+bD + b is one-to-one, so a careful text always says which space. The book, for instance, works with a finite-dimensional representation of (D+b)−1(D+b)^{-1} in Section 5.2.8, on a space closed under differentiation.

Watch the exponents. The nn-th derivative is written f(n)f^{(n)}, with parentheses, and f(0)=ff^{(0)} = f. Without parentheses, fnf^n usually means a power, and f−1f^{-1} means the inverse function, not 1/f1/f. Trigonometry adds an inconsistency you simply have to memorize: sin⁡2x=(sin⁡x)2\sin^2 x = (\sin x)^2, but sin⁡−1x=arcsin⁡x\sin^{-1} x = \arcsin x.

Partial derivatives, ∂f/∂x\partial f/\partial x or fxf_x or ∂xf\partial_x f, differentiate in one variable with all others held fixed. The phrase "held fixed" carries information: when variables are related, you need to know which ones are being held fixed, and a careful text says so.

Integrals: bounds, dummies and parameters

In ∫abf(x,t) dx\int_a^b f(x,t)\,dx, the variable xx is bound by the integral and the dxdx tells you so. The result is a function of aa, bb and the parameter tt, but not of xx. A standard example makes the dependence explicit:

F(t)=∫01xt dx=1t+1,t>−1F(t) = \int_0^1 x^{t}\,dx = \frac{1}{t+1}, \qquad t > -1

Differentiating with respect to the parameter, under conditions that justify moving ddt\frac{d}{dt} inside the integral, gives F′(t)=∫01xtln⁡x dx=−1/(t+1)2F'(t) = \int_0^1 x^{t}\ln x\,dx = -1/(t+1)^2, and in general F(n)(t)=∫01xt(ln⁡x)n dx=(−1)nn!/(t+1)n+1F^{(n)}(t) = \int_0^1 x^{t}(\ln x)^n\,dx = (-1)^n n!/(t+1)^{n+1}. Using definite integrals to organize higher-order derivatives in a parameter is the starting theme of Chapter 1 of the book; the article The structure behind higher-order derivatives discusses it.

When the variable of integration would clash with a limit, rename the dummy. Write G(x)=∫0xf(u) duG(x) = \int_0^x f(u)\,du, not ∫0xf(x) dx\int_0^x f(x)\,dx; the second form uses one letter for two different roles.

Transform notation

The Laplace transform takes a function ff and produces a new function of ss:

L{f}(s)=∫0∞e−stf(t) dt\mathcal{L}\{f\}(s) = \int_0^\infty e^{-st} f(t)\,dt

Here tt is a dummy variable and ss is free. Texts often write L{f(t)}\mathcal{L}\{f(t)\} or L{t2}=2/s3\mathcal{L}\{t^2\} = 2/s^3; this is a harmless abuse of notation as long as you remember that the transform acts on the whole function, not on a value. Capital letters are a common convention, F=L{f}F = \mathcal{L}\{f\}, and L−1\mathcal{L}^{-1} denotes the inverse transform. A transform formula is incomplete without its region of convergence: L{t2}(s)=2/s3\mathcal{L}\{t^2\}(s) = 2/s^3 holds for Re⁡s>0\operatorname{Re} s > 0, and the integral diverges for real s≤0s \le 0.

Convolution is written (f∗g)(t)(f \ast g)(t). In the Laplace setting it means ∫0tf(τ) g(t−τ) dτ\int_0^t f(\tau)\,g(t-\tau)\,d\tau; in the Fourier setting the integral runs over the whole real line. Same symbol, different definition: check which one a text uses.

Matrices, vectors and indices

A matrix A=(aij)∈Rm×nA = (a_{ij}) \in \mathbb{R}^{m\times n} has mm rows and nn columns, and aija_{ij} is the entry in row ii, column jj: row first, always. Vectors are columns unless stated otherwise, ATA^{T} is the transpose, and the product is defined entrywise by (Ax)i=∑j=1naijxj(Ax)_i = \sum_{j=1}^{n} a_{ij}x_j. The Kronecker delta δij\delta_{ij} equals 11 when i=ji = j and 00 otherwise, so the identity matrix is I=(δij)I = (\delta_{ij}).

The matrix of differentiation on span{sin x, cos x}

Write DD as a matrix on the space spanned by sin⁡x\sin x and cos⁡x\cos x, using the ordered basis (sin⁡x,cos⁡x)(\sin x, \cos x).

  1. Apply DD to each basis function: Dsin⁡x=cos⁡x=0⋅sin⁡x+1⋅cos⁡xD\sin x = \cos x = 0\cdot\sin x + 1\cdot\cos x, and Dcos⁡x=−sin⁡x=−1⋅sin⁡x+0⋅cos⁡xD\cos x = -\sin x = -1\cdot\sin x + 0\cdot\cos x.
  2. The coordinates of these images are (0,1)(0, 1) and (−1,0)(-1, 0). By convention they become the columns of the matrix.
  3. Check on a general element f=αsin⁡x+βcos⁡xf = \alpha\sin x + \beta\cos x: Df=−βsin⁡x+αcos⁡xDf = -\beta\sin x + \alpha\cos x, with coordinates (−β,α)(-\beta, \alpha), which is exactly the matrix times (α,β)T(\alpha, \beta)^{T}.
[D]=(0−110)[D] = \begin{pmatrix} 0 & -1 \\ 1 & 0 \end{pmatrix}

Matrices of this kind, differentiation written as linear algebra on a space closed under differentiation, are the subject of Chapter 5 of the book and of the article From function spaces to matrix representation.

Summation and silent conventions

The general Leibniz rule for the nn-th derivative of a product shows several conventions at once:

(fg)(n)=∑k=0n(nk)f(k) g(n−k)(fg)^{(n)} = \sum_{k=0}^{n} \binom{n}{k} f^{(k)}\,g^{(n-k)}

The index kk is bound and nn is free. The terms k=0k = 0 and k=nk = n use f(0)=ff^{(0)} = f and 0!=10! = 1. For n=2n = 2 the sum expands to f′′g+2f′g′+fg′′f''g + 2f'g' + fg'', which you can confirm by differentiating fgfg twice. Other silent conventions worth knowing: an empty sum equals 00, an empty product equals 11, and in some physics and geometry texts a repeated index is summed automatically (the Einstein convention), so aijxja_{ij}x_j already means ∑jaijxj\sum_j a_{ij}x_j.

Glossary

NotationRead asWatch for
dydx\frac{dy}{dx}, y′y', DyDyderivative of yy with respect to xxDD acts on everything to its right
f(n)f^{(n)}nn-th derivative of ffnot the power fnf^n; f(0)=ff^{(0)} = f
sin⁡2x\sin^2 x, sin⁡−1x\sin^{-1} x(sin⁡x)2(\sin x)^2; arcsin⁡x\arcsin xthe exponent convention is not consistent
∂f/∂x\partial f/\partial xpartial derivative in xxwhich variables are held fixed
∫abf(x,t) dx\int_a^b f(x,t)\,dxintegral over xx; a function of aa, bb, ttxx is a dummy variable
L{f}(s)\mathcal{L}\{f\}(s)Laplace transform of ff, evaluated at ssthe region of convergence
(f∗g)(t)(f \ast g)(t)convolution of f and gLaplace version integrates over [0,t][0, t]
aija_{ij}entry in row ii, column jjrow index first
δij\delta_{ij}Kronecker delta11 if i=ji = j, otherwise 00
∑k=0n\sum_{k=0}^{n}sum over kk from 00 to nnkk is bound; the result depends on nn
(D+b)−1(D+b)^{-1}inverse of the operator D+bD + bneeds a stated function space
Common notation in university analysis and where readers most often go wrong.

References

  1. Kevin Houston. How to Think Like a Mathematician. Cambridge University Press, 2009. Includes guidance on reading and writing mathematical notation..
  2. Lokenath Debnath and Dambaru Bhatta. Integral Transforms and Their Applications, 3rd ed.. CRC Press, 2014. Transform and convolution conventions..