Machine Learning Quizzes
Jerry Xiao Two

Quizzes for Maching Learning Course CS405 and CS329 at SUSTech. (Before 2024 Fall Semester)

Quiz 1

Question

y=Ax+vy=Ax + v, where vv is a Gaussian noise.

  1. What is the optimal solution for xx?
  2. What is the optimal solution for xx if v∼N(0,R)v \sim \mathcal{N}(0, R)?
  3. What is the optimal solution for xx if v∼N(0,R)v \sim \mathcal{N}(0, R) and X∼N(0,aI)X \sim N(0, aI)?
  4. AA and XX are unknown, what is the optimal solution for xx?

Answer

  1. J(x)=12(y−Ax)T(y−Ax)J(x) = \frac{1}{2} (y - Ax)^T(y-Ax), ∂J∂x=0\frac{\partial J}{\partial x} = 0, x=(ATA)−1ATyx = (A^TA)^{-1}A^Ty
  2. J(x)=12(y−Ax)TR−1(y−Ax)J(x) = \frac{1}{2} (y - Ax)^TR^{-1}(y-Ax), ∂J∂x=0\frac{\partial J}{\partial x} = 0, x=(ATR−1A)−1ATR−1yx = (A^TR^{-1}A)^{-1}A^TR^{-1}y
  3. J(x)=12(y−Ax)TR−1(y−Ax)+12xT(aI)−1xJ(x) = \frac{1}{2} (y - Ax)^TR^{-1}(y-Ax) + \frac{1}{2} x^T(aI)^{-1}x, ∂J∂x=0\frac{\partial J}{\partial x} = 0, x=(ATR−1A+aI)−1ATR−1yx = (A^TR^{-1}A + aI)^{-1}A^TR^{-1}y
  4. We can distinguish two cases:
    1. For x: J(x)=12(y−Ax)TR−1(y−Ax)+12xT(aI)−1xJ(x) = \frac{1}{2} (y - Ax)^TR^{-1}(y-Ax) + \frac{1}{2} x^T(aI)^{-1}x, ∂J∂x=0\frac{\partial J}{\partial x} = 0, x=(ATR−1A+aI)−1ATR−1yx = (A^TR^{-1}A + aI)^{-1}A^TR^{-1}y
    2. For A: YT=XTATY^T = X^TA^T J(A)=12(Y−XA)TR−1(Y−XA)J(A) = \frac{1}{2} (Y - XA)^TR^{-1}(Y-XA), ∂J∂A=0\frac{\partial J}{\partial A} = 0, AT=(XR−1XT)−1XR−1YTA^T = (XR^{-1}X^T)^{-1}XR^{-1}Y^T

Quiz 2

Question

Y=AX+ωY = AX + \omega, where ω∼N(0,Q)\omega \sim \mathcal{N}(0, Q) and X∼N(μ0,Σ0)X \sim \mathcal{N}(\mu_0, \Sigma_0)

  1. What is p(Y∣X)p(Y|X)?
  2. What is p(Y)p(Y)?
  3. What is p(X∣Y)p(X|Y)?
  4. What is p(Y′∣Y)p(Y'|Y)?

Answer

  1. p(Y∣X)∼N(AX,Q)p(Y|X) \sim \mathcal{N}(AX, Q) We regard XX as a constant under conditional probability.
  2. p(Y)∼∫p(Y∣X)p(X)dx∼N(Aμ0,AΣ0AT+Q)p(Y) \sim \int p(Y|X) p(X) d x \sim \mathcal{N}(A\mu_0, A\Sigma_0 A^T + Q).

    var[Y]=var[AX]+var[ω]=AΣ0AT+Qvar[Y] = var[AX] + var[\omega] = A\Sigma_0 A^T + Q

  3. Assume that p(X∣Y)∼N(m,L)p(X|Y) \sim \mathcal{N}(m, L), then we can use the equality of quadratic from to solve the problems.
    1. p(X∣Y)∼p(Y∣X)p(X)=N(Y∣AX,Q)N(X∣μ0,Σ9)p(X|Y) \sim p(Y|X)p(X) = \mathcal{N}(Y| AX, Q) \mathcal{N}(X| \mu_0, \Sigma_9)
    2. −12(x−m)TL−1(x−m)∝−12(y−Ax)TQ−1(y−Ax)−12(x−μ0)TΣ0−1(x−μ0)-\frac{1}{2}(x-m)^T L^{-1}(x-m) \propto -\frac{1}{2}(y-Ax)^T Q^{-1}(y-Ax) -\frac{1}{2}(x-\mu_0)^T \Sigma_0^{-1}(x-\mu_0)
    3. We can get the result:

    L−1=ATQ−1A+Σ0−1L−1m=ATQ−1y+Σ0−1μ0\begin{aligned} L^{-1} &= A^T Q^{-1} A + \Sigma_0^{-1} \\ L^{-1} m &= A^T Q^{-1} y + \Sigma_0^{-1} \mu_0 \end{aligned}

    1. By applying [A+BCD]−1=A−1−A−1B[C−1+DA−1B]−1DA−1[A+BCD]^{-1} = A^{-1} - A^{-1}B[C^{-1} + DA^{-1}B]^{-1} D A^{-1}

    L=(I−KA)Σ0m=μ0+K(y−Aμ0)\begin{aligned} L &= (I - KA) \Sigma_0 \\ m &= \mu_0 + K(y - A\mu_0) \end{aligned}

    where K=Σ0AT(ATΣ0A+Q)−1K = \Sigma_0 A^T (A^T\Sigma_0A + Q)^{-1}
  4. p(Y′∣Y)∼∫p(Y′∣X)p(X∣Y)dx∼N(Am,ALAT+Q)p(Y'|Y) \sim \int p(Y'|X) p(X|Y) d x \sim \mathcal{N}(Am, AL A^T + Q). The same format as question 2.

Quiz 3

  1. Learning: p(θ∣D)∝p(D∣θ)p(θ)p(\theta|\mathcal{D}) \propto p(\mathcal{D}|\theta)p(\theta)
  2. Prediction: p(Dnew∣D)=∫p(Dnew∣θ)p(θ∣D)dθp(\mathcal{D}^{new}|\mathcal{D}) = \int p(\mathcal{D}^{new}|\theta)p(\theta|\mathcal{D})d\theta
  3. Evaluation: p(D)=∫p(D∣θ)p(θ)dθp(\mathcal{D}) = \int p(\mathcal{D}|\theta)p(\theta)d\theta

Question

Given t=Φ(x)ω+vt = \Phi(x) \omega + v where Φ(x)=[1,x,x...,xM]\Phi(x) = [1, x, x..., x^M] and v∼N(0,β−1)v \sim \mathcal{N}(0, \beta^{-1}), D={[x1,...,xN],[t1,...,tN]}\mathcal{D} = \{[x_1,...,x_N], [t_1, ..., t_N]\}

  1. What is the solution of ωML\omega_{ML}?
  2. What is the solution of ωMAP\omega_{MAP} if ω∼N(0,α−1I)\omega \sim \mathcal{N}(0, \alpha^{-1}I)?
  3. What is the predictive distribution if Dnew={xnew,tnew}\mathcal{D}^{new} = \{x^{new}, t^{new}\}?
  4. What is the model evaluation?

Answer

  1. J(ω)=β2(T−Φω)T(T−Φω)→ωML=(ΦTΦ)−1ΦTTJ(\omega) = \frac{\beta}{2}(T-\Phi \omega)^T(T-\Phi\omega) \rightarrow \omega_{ML} = (\Phi^T\Phi)^{-1}\Phi^T T
  2. J(ω)=β2(T−Φω)T(T−Φω)+α2ωTω→ωMAP=(βΦTΦ+αI)−1βΦTTJ(\omega) = \frac{\beta}{2}(T-\Phi\omega)^T(T-\Phi\omega) + \frac{\alpha}{2} \omega^T\omega \rightarrow \omega_{MAP} = (\beta\Phi^T\Phi + \alpha I)^{-1}\beta\Phi^T T
  3. N(Φ(xnew)ωMAP,Φ(xnew)ΣMAPΦ(xnew)T+βI)\mathcal{N}(\Phi(x^{new})\omega_{MAP}, \Phi(x^{new})\Sigma_{MAP}\Phi(x^{new})^T+\beta I)
  4. N(0,α−1ΦΦT+β−1I)\mathcal{N}(0, \alpha^{-1}\Phi\Phi^T+\beta^{-1}I)

Quiz 4

Question

For y=σ(Φ(x)w)y = \sigma(\Phi(x) w), and D={[x1,...,xN],[t1,...,tN]}\mathcal{D} = \{[x_1,...,x_N], [t_1, ..., t_N]\}, where σ(x)=11+e−x\sigma(x) = \frac{1}{1+e^{-x}}.

  1. What is the solution of wMLw_{ML}?
  2. What is the solution of wMAPw_{MAP} if w∼N(m0,S0)w \sim \mathcal{N}(m_0, S_0)?
  3. What is the predictive distribution if Dnew={xnew,tnew=1}\mathcal{D}^{new} = \{x^{new}, t^{new}=1\}?
  4. What is the model evaluation?

Answer

  1. J(w)=−∑n=1N{tnlogyn+(1−tn)log(1−yn)} b=▽J(w)=∑n=1NϕT(yn−tn) H=▽▽J(w)=∑n=1Nyn(1−yn)ϕnTϕJ(w) = -\sum_{n=1}^N \{t_n \log y_n + (1-t_n) \log(1-y_n)\} \ b = \triangledown J(w) = \sum_{n=1}^N \phi^T(y_n - t_n) \ H = \triangledown \triangledown J(w) = \sum_{n=1}^N y_n(1-y_n) \phi_n^T\phi

    Because σ\sigma is not a linear function, there are no explicit solution to find wMLw_{ML}. We can use the gradient descent method to find the solution.

    w+→w−H−1bw^+ \rightarrow w - H^{-1}b

  2. J(w)=−∑n=1N{tnlogyn+(1−tn)log(1−yn)}+12(w−m0)TS0−1(w−m0)J(w) = -\sum_{n=1}^N \{t_n \log y_n + (1-t_n) \log(1-y_n)\} + \frac{1}{2}(w-m_0)^TS_0^{-1}(w-m_0)

    Therefore,and H=▽▽J(w)=∑n=1Nyn(1−yn)ϕnTϕ+S0−1H = \triangledown \triangledown J(w) =\sum_{n=1}^N y_n(1-y_n) \phi_n^T\phi + S_0^{-1}

  3. p(tnew=1∣xnew,D)=∫p(tnew=1∣w)p(w∣D)dw=∫σ(ϕneww)N(wMAP,H−1)dwp(t^{new}=1 | x^{new}, \mathcal{D}) = \int p(t^{new}=1|w)p(w|\mathcal{D})dw = \int \sigma(\phi^{new} w) \mathcal{N}(w_{MAP}, H^{-1})dw

    σ(κ(σa2)μa)\sigma(\kappa(\sigma_a^2)\mu_a)

  4. ∑n=1N[tnlnyn+(1−tn)ln(1−yn)]MAP−12(wMAP−m0)TS0−1(wMAP−m0)+M2ln2π−12ln∣H∣MAP\sum_{n=1}^N \left[t_n \ln y_n + (1 - t_n) \ln (1 - y_n)\right]_{\text{MAP}} - \frac{1}{2} (w_\text{MAP}- m_0)^T S_0^{-1}(w_\text{MAP}- m_0)+ \frac{M}{2} \ln 2\pi - \frac{1}{2}\ln |H|_\text{MAP}

Powered by Hexo & Theme Keep
This site is deployed on