# Entropy-Rate Selection for Partially Observed Processes
**Authors**:
- Oleg Kiriukhin (City University of Hong Kong)
(April 2026)
## Abstract
I formulate an entropy-rate maximization problem at the observable level for stochastic processes observed through an information-reducing observation map. For a visible stationary law, the map determines an observational fiber of hidden stationary laws generating that law. In the finite-state finite-memory setting, retained visible constraints determine a feasible class of stationary $(r+1)$ -block laws, and the entropy maximizer is defined as the entropy-rate maximizer on this class.
The paper formulates entropy-rate maximization on feasible classes induced by partial observability and develops a structural theory for the resulting maximizer. I prove existence and uniqueness of the maximizer, with uniqueness under a fixed-context-marginal hypothesis and, more generally, via a strict-concavity characterization by row proportionality. Two global characterization regimes are central: a fixed one-point marginal yields the i.i.d. maximizer, and a fixed $r$ -block law yields the $(r-1)$ -step Markov extension. The gap functional equals a conditional mutual information and vanishes exactly at the maximizing completion. I also derive optimality conditions, local geometry of the maximizer, a latent random-mapping realization that leaves the visible law unchanged, and a local empirical consistency theorem, and illustrate the framework by an aliased hidden-state example.
Keywords. entropy rate, partial observability, observational fibers, entropy maximization, stationary block laws, finite-state Markov processes, Markov extension, conditional mutual information, hidden stochastic processes.
## 1 Introduction
Stochastic models are often underidentified by the information structure through which they are observed. Distinct hidden mechanisms may generate the same visible law, so one faces an observational equivalence class rather than a uniquely recoverable latent model. I take the observation experiment, together with the retained visible observables, as primitive and ask whether they determine a preferred visible completion within a finite-dimensional block-law class. This perspective is motivated by Blackwell’s comparison of experiments, where the primitive object is the information structure rather than an externally chosen parametric family [1, 2].
Entropy maximization under constraints and entropy calculations for hidden processes are classical themes. The novelty here lies in formulating an entropy-rate maximization problem at the observable level on feasible classes induced by partial observability, and in developing the corresponding structural theory of the maximizer. The emphasis is therefore on the visible law determined by the experiment rather than on selecting a hidden model from an observationally equivalent family.
For a visible law, I define the observational fiber as the class of hidden stationary laws generating that law under a fixed observation experiment. Restricting to finite alphabets and finite memory, I work with stationary $(r+1)$ -block laws on the visible alphabet. In block-law coordinates, stationarity is a finite-dimensional linear constraint and the feasible set is a convex subset of the simplex on $A^r+1$ . An entropy maximizer is an entropy-rate maximizer on the resulting feasible class, and it selects the visible completion with maximal residual uncertainty, hence minimal serial organization beyond what the retained observables force.
Two global characterization regimes are central. If the retained observable fixes the one-point marginal, the maximizer is the i.i.d. process with that marginal. If it fixes the full stationary $r$ -block law, the maximizer is the $(r-1)$ -step Markov extension. In the latter regime, the gap functional is a conditional mutual information and vanishes exactly at the maximizing completion.
The main results are as follows. I define observational fibers and the corresponding feasible classes determined by the retained observables at the visible block-law level, and prove existence and uniqueness of the entropy maximizer, with uniqueness under a fixed-context-marginal hypothesis and, more generally, via a strict-concavity characterization by row proportionality. I then prove two global characterization theorems: fixed one-point marginals yield the i.i.d. maximizer, and fixed $r$ -block laws yield the $(r-1)$ -step Markov maximizer. I also identify the gap functional with conditional mutual information, which vanishes exactly at the maximizing completion.
At the local level, I derive optimality conditions and a kernel formula, identify the intrinsic tangent space and restricted Hessian on a fixed-support face, prove a strict-concavity criterion on affine slices of a common positive-support face, and obtain local geometry together with a quadratic expansion of the gap functional. I also prove that any selected visible law admits a latent random-mapping realization that leaves the visible law unchanged, establish a local empirical consistency theorem under a full-rank moment-map hypothesis, and include an aliased hidden-state example showing that a maximizing visible completion need not resolve hidden underidentification.
## 2 Observational Fibers and Entropy Maximization
Let $H$ and $A$ be finite alphabets, and let
$$
Π:P_stat(H^ℤ)→P_stat(A^ℤ), Q↦Π_\#Q \tag{1}
$$
be a fixed measurable observation map on stationary path laws. For a visible stationary law
$$
ν∈P_stat(A^ℤ),
$$
define the observational fiber
$$
E_Π(ν):=\{Q∈P_stat(H^ℤ):Π_\#Q=ν\}. \tag{2}
$$
To obtain a finite-dimensional selector, fix a memory length $r≥ 1$ and work with stationary $(r+1)$ -block laws on the visible alphabet. Let
$$
u(c,a), c∈ A^r, a∈ A,
$$
denote a probability distribution on $A^r+1$ .
**Definition 2.1 (Context marginal)**
*For a block law $u$ on $A^r+1$ , define its context marginal by
$$
η_u(c):=∑_a∈ Au(c,a), c∈ A^r.
$$*
**Definition 2.2 (Stationary consistency)**
*A block law $u$ on $A^r+1$ is called stationary-consistent if its left and right $r$ -block marginals agree, that is,
$$
∑_α∈ Au(α,c_1,\dots,c_r)=∑_β∈ Au(c_1,\dots,c_r,β)
$$
for every $(c_1,\dots,c_r)∈ A^r$ .*
Fix retained observable features
$$
G_1,\dots,G_m:A^r+1→ℝ,
$$
and, for each visible stationary law $ν$ , define the associated target vector
$$
b(ν):=\bigl(b_1(ν),\dots,b_m(ν)\bigr)∈ℝ^m.
$$
In the finite-state setting, the induced visible feasible class is
$$
U_Π(ν):=\Bigl\{u∈Δ(A^r+1):u is stationary-consistent and ∑_c,au(c,a)G_j(c,a)=b_j(ν), j=1,\dots,m\Bigr\}, \tag{3}
$$
Whenever $η_u(c)>0$ on the active support, define the induced conditional kernel by
$$
p_u(a\mid c):=\frac{u(c,a)}{η_u(c)}. \tag{4}
$$
The entropy-rate functional is then
$$
J(u):=-∑_c∈ A^r∑_a∈ Au(c,a)\log\frac{u(c,a)}{η_u(c)}. \tag{5}
$$
Equivalently,
$$
J(u)=∑_c∈ A^rη_u(c) H\bigl(p_u(·\mid c)\bigr). \tag{6}
$$
For stationary finite-state Markov processes, this is the entropy rate
$$
H(X_r\mid X_0^r-1).
$$
**Definition 2.3 (Entropy maximizer)**
*An entropy maximizer is any maximizer
$$
u^⋆∈\arg\max_u∈U_Π(ν)J(u).
$$
Whenever $η_u^⋆(c)>0$ , the induced active-support kernel is
$$
p^⋆(a\mid c):=\frac{u^⋆(c,a)}{η_u^⋆(c)}.
$$*
**Assumption 2.4 (Global standing assumptions)**
*1. The retained observable constraints are finitely many linear equalities in the block-law coordinates $u(c,a)$ .
1. The feasible set $U_Π(ν)$ is nonempty.
1. For the global uniqueness theorem, the feasible set fixes the context marginal: there exists $\bar{η}$ such that $η_u=\bar{η}$ for all $u∈U_Π(ν)$ .*
**Assumption 2.5 (Local regularity assumptions)**
*1. The selected point $u^⋆$ lies in the relative interior of a fixed-support face.
1. A smooth local chart $(b,ξ)↦ u(b,ξ)$ exists near $u^⋆$ , where $b$ are retained observable coordinates and $ξ$ are fiber coordinates.*
On the positive-support region, block-law and kernel parametrizations are equivalent:
$$
u(c,a)=η_u(c)p_u(a\mid c). \tag{7}
$$
## 3 Existence, Uniqueness, and Optimality Conditions
**Theorem 3.1 (Existence)**
*Under Assumptions (A1)–(A2), the optimization problem
$$
\max_u∈U_Π(ν)J(u)
$$
admits at least one maximizer.*
* Proof*
The variable $u$ lies in the probability simplex on the finite set $A^r+1$ . The stationary-consistency equations and the experiment-induced restrictions are linear in $u(c,a)$ , so the feasible set $U_Π(ν)$ is a closed subset of a finite-dimensional simplex. Hence $U_Π(ν)$ is compact and convex. The map $x↦-x\log x$ extends continuously to $[0,1]$ by the convention $0\log 0=0$ . Therefore the function
$$
u↦-∑_c,au(c,a)\log u(c,a)
$$
is continuous on the simplex. Since
$$
J(u)=-∑_c,au(c,a)\log u(c,a)+∑_cη_u(c)\logη_u(c),
$$
and the marginal map $u↦η_u$ is linear, $J$ is continuous on $U_Π(ν)$ . By compactness, $J$ attains its maximum on $U_Π(ν)$ . ∎
**Theorem 3.2 (Uniqueness under fixed context marginals)**
*Assume the hypotheses of theorem 3.1, and assume in addition that there exists a probability vector $\bar{η}$ on $A^r$ such that $η_u=\bar{η}$ for every $u∈U_Π(ν)$ . Then the maximizer of $J$ over $U_Π(ν)$ is unique.*
* Proof*
Assumption (A3) states that the feasible set fixes the context marginal: there exists $\bar{η}$ such that $η_u=\bar{η}$ for every feasible $u$ . Hence on $U_Π(ν)$ one has
$$
J(u)=-∑_c,au(c,a)\log u(c,a)+∑_c\bar{η}(c)\log\bar{η}(c).
$$
The second term is constant over the feasible set, so maximizing $J$ is equivalent to maximizing the Shannon entropy of the block law itself:
$$
H(u):=-∑_c,au(c,a)\log u(c,a).
$$
Shannon entropy is strictly concave on a simplex [3]. Therefore $H(u)$ is strictly concave on the affine hull of the feasible set, and hence on the convex set $U_Π(ν)$ . A strictly concave function has at most one maximizer on a convex set. Existence was proved in theorem 3.1, so the maximizer is unique. ∎
**Proposition 3.3 (Concavity and equality criterion)**
*Let $u$ and $v$ be stationary $(r+1)$ -block laws on $A^r+1$ , and let $t∈(0,1)$ . Then
$$
J(tu+(1-t)v)≥ tJ(u)+(1-t)J(v).
$$
Moreover, equality holds if and only if for every context $c∈ A^r$ the rows $u(c,·)$ and $v(c,·)$ are proportional. Equivalently, equality holds if and only if
$$
η_v(c)u(c,a)=η_u(c)v(c,a) for every c∈ A^r, a∈ A.
$$
On any common positive-support face, this is further equivalent to equality of the conditional kernels in every active context.*
* Proof*
Write
$$
w:=tu+(1-t)v.
$$
For each context $c∈ A^r$ , define the row masses
$$
η_u(c)=∑_au(c,a), η_v(c)=∑_av(c,a), η_w(c)=tη_u(c)+(1-t)η_v(c).
$$
Then
$$
J(u)=∑_cη_u(c)H\bigl(p_u(·\mid c)\bigr), J(v)=∑_cη_v(c)H\bigl(p_v(·\mid c)\bigr), J(w)=∑_cη_w(c)H\bigl(p_w(·\mid c)\bigr),
$$
where the row conditional laws are defined whenever the corresponding row mass is positive. For a fixed context $c$ with $η_w(c)>0$ , define
$$
α_c:=\frac{tη_u(c)}{η_w(c)}, 1-α_c:=\frac{(1-t)η_v(c)}{η_w(c)}.
$$
Then $α_c∈[0,1]$ and, for each symbol $a∈ A$ ,
$$
p_w(a\mid c)=α_cp_u(a\mid c)+(1-α_c)p_v(a\mid c)
$$
whenever both row masses are positive. By concavity of Shannon entropy on the simplex [3],
$$
H\bigl(p_w(·\mid c)\bigr)≥α_cH\bigl(p_u(·\mid c)\bigr)+(1-α_c)H\bigl(p_v(·\mid c)\bigr).
$$
Multiplying by $η_w(c)$ yields
$$
η_w(c)H\bigl(p_w(·\mid c)\bigr)≥ tη_u(c)H\bigl(p_u(·\mid c)\bigr)+(1-t)η_v(c)H\bigl(p_v(·\mid c)\bigr).
$$
Summing over $c∈ A^r$ gives the asserted concavity inequality. For the equality statement, equality in the entropy concavity inequality for a fixed context $c$ holds if and only if the two conditional distributions in that row coincide whenever both mixture weights are positive. Equivalently, whenever $η_u(c),η_v(c)>0$ ,
$$
p_u(a\mid c)=p_v(a\mid c) for every a∈ A.
$$
Multiplying through by the row masses gives
$$
η_v(c)u(c,a)=η_u(c)v(c,a) for every a∈ A.
$$
If one of the row masses is zero, equality can hold only when the other row is also zero, which is again exactly the same proportionality condition. Therefore equality in the global concavity inequality holds if and only if the displayed proportionality identities hold in every context. These identities are equivalent to rowwise proportionality of $u(c,·)$ and $v(c,·)$ . Finally, on a common positive-support face all active row masses are positive, so dividing by $η_u(c)η_v(c)$ shows that the same condition is equivalent to
$$
p_u(·\mid c)=p_v(·\mid c) for every active context c.
$$ ∎
**Theorem 3.4 (Strict-concavity characterization by row proportionality)**
*Let $K$ be a convex set of stationary $(r+1)$ -block laws on $A^r+1$ . Then the following are equivalent:
1. $J$ is strictly concave on $K$ .
1. For every distinct pair $u,v∈ K$ , there exist a context $c∈ A^r$ and a symbol $a∈ A$ such that
$$
η_v(c)u(c,a)≠η_u(c)v(c,a).
$$
1. No two distinct points of $K$ are rowwise proportional in every context.
If, in addition, every element of $K$ belongs to a common positive-support face, then these conditions are also equivalent to:
1. For every distinct pair $u,v∈ K$ , there exists an active context $c$ such that
$$
p_u(·\mid c)≠ p_v(·\mid c).
$$*
* Proof*
The equivalence of (ii) and (iii) is the coordinate form of rowwise proportionality. If all elements of $K$ belong to a common positive-support face, then for each active context $c$ every row mass is positive, and dividing the identity
$$
η_v(c)u(c,a)=η_u(c)v(c,a)
$$
by $η_u(c)η_v(c)$ shows that (ii) is equivalent to (iv). Assume (i). If (ii) failed, then there would exist distinct $u,v∈ K$ such that
$$
η_v(c)u(c,a)=η_u(c)v(c,a) for every c∈ A^r, a∈ A.
$$
By proposition 3.3, equality would then hold along the whole segment joining $u$ and $v$ :
$$
J((1-t)u+tv)=(1-t)J(u)+tJ(v) for every t∈(0,1),
$$
which contradicts strict concavity on $K$ . Thus (i) implies (ii). Assume (ii). Let $u,v∈ K$ be distinct and let $t∈(0,1)$ . Since $K$ is convex, the point
$$
w:=(1-t)u+tv
$$
belongs to $K$ . By proposition 3.3,
$$
J(w)≥(1-t)J(u)+tJ(v),
$$
and equality holds if and only if $u$ and $v$ are rowwise proportional in every context. Hypothesis (ii) rules out that equality case for distinct $u$ and $v$ . Therefore
$$
J((1-t)u+tv)>(1-t)J(u)+tJ(v) for all distinct u,v∈ K, t∈(0,1),
$$
which is exactly strict concavity of $J$ on $K$ . Thus (ii) implies (i). This proves the equivalence of (i)–(iii), and also of (i)–(iv) on a common positive-support face. ∎
**Corollary 3.5 (Global uniqueness on the feasible class beyond (A3))**
*Assume (A1)–(A2). Suppose that for every distinct feasible pair $u,v∈U_Π(ν)$ there exist a context $c∈ A^r$ and a symbol $a∈ A$ such that
$$
η_v(c)u(c,a)≠η_u(c)v(c,a).
$$
Then the maximizer of $J$ over $U_Π(ν)$ is unique.*
* Proof*
By theorem 3.1, the feasible class $U_Π(ν)$ is nonempty, compact, and convex, and $J$ attains a maximum on it. The stated hypothesis implies, via theorem 3.4, that $J$ is strictly concave on $U_Π(ν)$ . A strictly concave function has at most one maximizer on a convex set, so the maximizer is unique. ∎
**Corollary 3.6 (Kernel-separation criterion on a common positive-support face)**
*Assume (A1)–(A2) and suppose every feasible block law in $U_Π(ν)$ belongs to a common positive-support face. Assume, in addition, that for every distinct feasible pair $u,v∈U_Π(ν)$ there exists an active context $c$ such that
$$
p_u(·\mid c)≠ p_v(·\mid c).
$$
Then the maximizer of $J$ over $U_Π(ν)$ is unique.*
* Proof*
Under the common positive-support assumption, theorem 3.4 shows that the kernel-separation hypothesis is equivalent to strict concavity of $J$ on $U_Π(ν)$ . The conclusion then follows from corollary 3.5. ∎
**Remark 3.7**
*theorem 3.4 is strictly broader than theorem 3.2. Assumption (A3) eliminates all variation of the context marginal across the feasible set, whereas the broader characterization excludes only the flat directions along which every row changes by a scalar factor while the conditional kernels remain unchanged.*
#### A sufficient condition for Assumption (A3).
Suppose that for each context $c∈ A^r$ , the context-indicator function
$$
I_c(c^\prime,a):=1\{c^\prime=c\}, (c^\prime,a)∈ A^r+1,
$$
belongs to the linear span of the retained observable features $G_1,\dots,G_m$ together with the constant function $1$ on $A^r+1$ . Then the feasible set $U_Π(ν)$ fixes the full context marginal, so Assumption (A3) holds and theorem 3.2 applies.
* Proof*
For any feasible block law $u∈U_Π(ν)$ ,
$$
η_u(c)=∑_a∈ Au(c,a)=∑_c^\prime,au(c^\prime,a)I_c(c^\prime,a).
$$
By hypothesis, one may write
$$
I_c=α_c,0+∑_j=1^mα_c,jG_j
$$
for suitable coefficients depending on $c$ . Therefore
$$
η_u(c)=α_c,0∑_c^\prime,au(c^\prime,a)+∑_j=1^mα_c,j∑_c^\prime,au(c^\prime,a)G_j(c^\prime,a).
$$
The first term equals $α_c,0$ by normalization, and each remaining term is fixed on $U_Π(ν)$ by the retained moment constraints. Hence $η_u(c)$ is the same for every feasible $u$ . Hence the full context marginal is fixed on $U_Π(ν)$ . ∎
## 4 Global Characterization Theorems
**Theorem 4.1 (Fixed marginal case)**
*Let $A$ be a finite alphabet, and let $π=(π_a)_a∈ A$ be a probability vector on $A$ . Consider the feasible class
$$
U(π):=≤ft\{u=(u_a,b)_a,b∈ A:u_a,b≥ 0,∑_b∈ Au_a,b=π_a,∑_a∈ Au_a,b=π_b\right\}.
$$
Then the entropy-rate functional
$$
J(u)=-∑_a,b∈ Au_a,b\log\frac{u_a,b}{π_a}
$$
has the unique maximizer
$$
u^⋆_a,b=π_aπ_b, a,b∈ A.
$$
That is, the entropy maximizer is the i.i.d. process with one-point law $π$ .*
* Proof*
Let $(X_t)_t∈ℤ$ be a stationary first-order process with two-block law $u$ . Because $u∈ U(π)$ , both $X_0$ and $X_1$ have marginal law $π$ . By the entropy-rate definition,
$$
J(u)=H(X_1\mid X_0).
$$
By monotonicity of conditional entropy, applied to the conditioning variables $X_0$ and the trivial sigma-field [3],
$$
H(X_1\mid X_0)≤ H(X_1)=H(π).
$$
Hence $J(u)≤ H(π)$ for every $u∈ U(π)$ . Now define $u^⋆$ by $u^⋆_a,b=π_aπ_b$ . Then $u^⋆∈ U(π)$ because
$$
∑_b∈ Au^⋆_a,b=π_a∑_b∈ Aπ_b=π_a, ∑_a∈ Au^⋆_a,b=π_b∑_a∈ Aπ_a=π_b.
$$
The corresponding process is i.i.d. with one-point law $π$ . Therefore $X_1$ is independent of $X_0$ , and thus
$$
J(u^⋆)=H(X_1\mid X_0)=H(X_1)=H(π).
$$
Hence $u^⋆$ attains the maximum. Suppose that $u∈ U(π)$ also satisfies $J(u)=H(π)$ . Then
$$
H(X_1\mid X_0)=H(X_1).
$$
Equality holds if and only if $X_1$ is independent of $X_0$ . Since both marginals equal $π$ , independence implies
$$
u_a,b=ℙ(X_0=a,X_1=b)=π_aπ_b, a,b∈ A.
$$
Hence $u=u^⋆$ . ∎
**Corollary 4.2 (Binary fixed-mean case)**
*Let $A=\{0,1\}$ , and let the retained observable be the stationary mean
$$
m=ℙ(X_t=1)∈[0,1].
$$
Then the unique entropy-rate maximizer among stationary binary first-order laws with mean $m$ is the i.i.d. Bernoulli $(m)$ law.*
* Proof*
Apply theorem 4.1 with $π=(1-m,m)$ . ∎
**Theorem 4.3 (Fixedrr-block case)**
*Let $r≥ 1$ , let $A$ be a finite alphabet, and let $μ$ be a probability law on $A^r$ . Adopt the convention that $A^0=\{∅\}$ and that the empty-block marginal equals $μ(∅)=1$ when $r=1$ . Consider the feasible class
$$
U(μ):=≤ft\{u=(u(c,a))_c∈ A^r, a∈ A:\begin{array}[]{l}u(c,a)≥ 0,\\[3.00003pt]
∑_a∈ Au(c,a)=μ(c) for every c∈ A^r,\\[3.00003pt]
u is stationary-consistent\end{array}\right\}.
$$
For $u∈ U(μ)$ , write
$$
J(u)=-∑_c∈ A^r∑_a∈ Au(c,a)\log\frac{u(c,a)}{μ(c)}=H(X_r\mid X_0^r-1),
$$
where $(X_t)$ is any stationary process with $(r+1)$ -block law $u$ . When $r=1$ , the expression $X_1^r-1$ below is interpreted as the trivial conditioning sigma-field, so
$$
H_μ(X_r\mid X_1^r-1)=H_μ(X_1).
$$ For $s=(s_1,\dots,s_r-1)∈ A^r-1$ with $μ(s)>0$ , define
$$
q(a\mid s):=\frac{μ(s,a)}{μ(s)}, a∈ A.
$$
For $s∈ A^r-1$ with $μ(s)=0$ , fix an arbitrary probability vector $q(·\mid s)$ on $A$ . Define
$$
u^⋆(c,a)=μ(c) q(a\mid c_2,\dots,c_r), c=(c_1,\dots,c_r)∈ A^r, a∈ A.
$$
Then $u^⋆∈ U(μ)$ and $u^⋆$ is the unique maximizer of $J$ on $U(μ)$ . That is, among all stationary $(r+1)$ -block laws extending $μ$ , the entropy-rate selector is the $(r-1)$ -step Markov extension.*
* Proof*
Let $u∈ U(μ)$ , and let $(X_t)_t∈ℤ$ be a stationary process with $(r+1)$ -block law $u$ . Because $∑_a∈ Au(c,a)=μ(c)$ for every $c∈ A^r$ , the law of $(X_0,\dots,X_r-1)$ is $μ$ . Because $u$ is stationary-consistent, the law of $(X_1,\dots,X_r)$ is also $μ$ . By definition,
$$
J(u)=H(X_r\mid X_0^r-1).
$$
Since $X_1^r-1$ is a function of $X_0^r-1$ , conditioning on the larger sigma-field cannot increase entropy [3], so
$$
H(X_r\mid X_0^r-1)≤ H(X_r\mid X_1^r-1).
$$
Moreover, the joint law of $(X_1,\dots,X_r)$ is fixed and equal to $μ$ , so the quantity on the right-hand side depends only on $μ$ . Therefore every feasible law satisfies
$$
J(u)≤ H_μ(X_r\mid X_1^r-1).
$$ To verify $u^⋆∈ U(μ)$ , note first that
$$
∑_a∈ Au^⋆(c,a)=μ(c)∑_a∈ Aq(a\mid c_2,\dots,c_r)=μ(c) c∈ A^r,
$$
so $u^⋆$ has the required context marginal. It remains to verify stationarity-consistency. Fix $d=(d_1,\dots,d_r)∈ A^r$ . Then
$$
∑_a∈ Au^⋆(a,d)=∑_a∈ Aμ(a,d_1,\dots,d_r-1) q(d_r\mid d_1,\dots,d_r-1).
$$
If $μ(d_1,\dots,d_r-1)>0$ , the factor $q(d_r\mid d_1,\dots,d_r-1)$ is equal to
$$
\frac{μ(d_1,\dots,d_r)}{μ(d_1,\dots,d_r-1)}.
$$
Hence
$$
∑_a∈ Au^⋆(a,d)=\Bigl(∑_a∈ Aμ(a,d_1,\dots,d_r-1)\Bigr)\frac{μ(d_1,\dots,d_r)}{μ(d_1,\dots,d_r-1)}.
$$
By marginalization of the probability law $μ$ ,
$$
∑_a∈ Aμ(a,d_1,\dots,d_r-1)=μ(d_1,\dots,d_r-1).
$$
Therefore
$$
∑_a∈ Au^⋆(a,d)=μ(d_1,\dots,d_r).
$$
If instead $μ(d_1,\dots,d_r-1)=0$ , then every term $μ(a,d_1,\dots,d_r-1)$ is zero, so
$$
∑_a∈ Au^⋆(a,d)=0.
$$
Also $μ(d_1,\dots,d_r)≤μ(d_1,\dots,d_r-1)=0$ , hence $μ(d_1,\dots,d_r)=0$ . Therefore in all cases
$$
∑_a∈ Au^⋆(a,d)=μ(d)=∑_a∈ Au^⋆(d,a),
$$
which is stationarity-consistency. Thus $u^⋆∈ U(μ)$ . Fix a context $c=(c_1,\dots,c_r)$ with $μ(c)>0$ . For every $a∈ A$ , the definition of $u^⋆$ and the identity $∑_b∈ Au^⋆(c,b)=μ(c)$ give
$$
ℙ_u^⋆(X_r=a\mid X_0^r-1=c)=\frac{u^⋆(c,a)}{μ(c)}=q(a\mid c_2,\dots,c_r).
$$
The right-hand side depends on $c$ only through $(c_2,\dots,c_r)$ , so
$$
ℙ_u^⋆(X_r=a\mid X_0^r-1)=ℙ_u^⋆(X_r=a\mid X_1^r-1) for every a∈ A,
$$
that is, $X_r⊥ X_0\mid X_1^r-1$ . Therefore
$$
H(X_r\mid X_0^r-1)=H(X_r\mid X_1^r-1).
$$
Since the law of $(X_1,\dots,X_r)$ is $μ$ ,
$$
J(u^⋆)=H_μ(X_r\mid X_1^r-1).
$$
Hence $u^⋆$ attains the maximum. Suppose that $u∈ U(μ)$ also attains the same maximal value. Then equality holds in the conditional-entropy inequality:
$$
H(X_r\mid X_0^r-1)=H(X_r\mid X_1^r-1).
$$
Since $X_1^r-1$ is a function of $X_0^r-1$ , the difference between these two entropies is the conditional mutual information of $X_0$ and $X_r$ given $X_1^r-1$ , namely
$$
H(X_r\mid X_1^r-1)-H(X_r\mid X_0^r-1)=I(X_0,X_r\mid X_1^r-1).
$$
Hence equality holds if and only if
$$
I(X_0,X_r\mid X_1^r-1)=0,
$$
that is, if and only if
$$
X_r⊥ X_0\mid X_1^r-1.
$$ Now fix $c=(c_1,\dots,c_r)∈ A^r$ with $μ(c)>0$ and let $a∈ A$ . By the chain rule and the conditional independence just established,
$$
u(c,a)=ℙ(X_0^r-1=c,X_r=a)=ℙ(X_0^r-1=c) ℙ(X_r=a\mid X_0^r-1=c)
$$
$$
=μ(c) ℙ(X_r=a\mid X_1^r-1=c_2,\dots,c_r)=μ(c) q(a\mid c_2,\dots,c_r)=u^⋆(c,a).
$$
Thus $u(c,a)=u^⋆(c,a)$ for every $c∈ A^r$ with $μ(c)>0$ and every $a∈ A$ . If $μ(c)=0$ , then feasibility gives $∑_a∈ Au(c,a)=0$ . Since every coordinate is nonnegative, this implies $u(c,a)=0$ for every $a∈ A$ . By the definition of $u^⋆$ one also has $u^⋆(c,a)=0$ for every $a∈ A$ . Therefore $u=u^⋆$ on all coordinates. ∎
**Corollary 4.4 (Gap and conditional mutual information)**
*Under the hypotheses of theorem 4.3, define
$$
Δ_μ(u):=H_μ(X_r\mid X_1^r-1)-J(u).
$$
Then, for every $u∈ U(μ)$ ,
$$
Δ_μ(u)=I_u(X_0,X_r\mid X_1^r-1)≥ 0.
$$
Moreover,
$$
Δ_μ(u)=0 \Longleftrightarrow u=u^⋆.
$$*
* Proof*
Let $u∈ U(μ)$ . Because the law of $(X_1,\dots,X_r)$ is fixed and equal to $μ$ , one has
$$
H_μ(X_r\mid X_1^r-1)=H_u(X_r\mid X_1^r-1).
$$
Using the definition of $J(u)$ from theorem 4.3,
$$
Δ_μ(u)=H_u(X_r\mid X_1^r-1)-H_u(X_r\mid X_0^r-1).
$$
Since $X_0^r-1=(X_0,X_1^r-1)$ , the defining identity for conditional mutual information gives
$$
I_u(X_0,X_r\mid X_1^r-1)=H_u(X_r\mid X_1^r-1)-H_u(X_r\mid X_0,X_1^r-1)=H_u(X_r\mid X_1^r-1)-H_u(X_r\mid X_0^r-1).
$$
Therefore
$$
Δ_μ(u)=I_u(X_0,X_r\mid X_1^r-1)≥ 0.
$$ Finally, $Δ_μ(u)=0$ if and only if $I_u(X_0,X_r\mid X_1^r-1)=0$ . This is equivalent to
$$
X_r⊥ X_0\mid X_1^r-1.
$$
By the uniqueness statement already proved in theorem 4.3, that conditional independence relation is equivalent to $u=u^⋆$ . Hence $Δ_μ(u)=0$ if and only if $u=u^⋆$ . ∎
The experiment-induced linear restrictions are
$$
∑_c,au(c,a)G_j(c,a)=b_j, j=1,\dots,m. \tag{8}
$$
The stationarity-consistency constraints can be written as
$$
∑_a∈ Au(c,a)-∑_a∈ Au(a,c)=0, c∈ A^r, \tag{9}
$$
where $u(a,c)$ denotes the block whose suffix of length $r$ is $c$ .
For a block $(c,a)∈ A^r+1$ , let $σ(c,a)$ denote its suffix of length $r$ .
**Proposition 4.5 (Optimality conditions)**
*Assume the unique maximizer $u^⋆$ lies in the relative interior of a fixed-support face of the feasible set. Then there exist multipliers $λ_1,\dots,λ_m$ , a scalar $γ$ , and stationarity multipliers $ψ(c)$ , $c∈ A^r$ , such that for each active coordinate $(c,a)$ ,
$$
\log\frac{u^⋆(c,a)}{η_u^⋆(c)}=-γ-∑_j=1^mλ_jG_j(c,a)-ψ(c)+ψ(σ(c,a)).
$$
Equivalently,
$$
u^⋆(c,a)=η_u^⋆(c)\exp\Bigl(-γ-∑_j=1^mλ_jG_j(c,a)-ψ(c)+ψ(σ(c,a))\Bigr).
$$*
* Proof*
Restrict the feasible set to the affine slice determined by the linear experiment restrictions, the normalization constraint, and the stationarity equations, and then to the fixed-support face on which $u^⋆(c,a)>0$ . On that relative interior, $J$ is $C^1$ , so the usual finite-dimensional Lagrange-multiplier conditions apply. Consider the Lagrangian
$$
L(u,λ,ψ,γ)=J(u)-∑_j=1^mλ_j\Bigl(∑_c,au(c,a)G_j(c,a)-b_j\Bigr)-γ\Bigl(∑_c,au(c,a)-1\Bigr)-∑_c∈ A^rψ(c)\Bigl(∑_au(c,a)-∑_au(a,c)\Bigr).
$$
Since
$$
J(u)=-∑_c,au(c,a)\log u(c,a)+∑_cη_u(c)\logη_u(c), η_u(c)=∑_au(c,a),
$$
its derivative with respect to an active coordinate $u(c,a)$ is
$$
\frac{∂}{∂ u(c,a)}\Bigl[-∑_c^\prime,a^{\prime}u(c^\prime,a^\prime)\log u(c^\prime,a^\prime)\Bigr]=-(\log u(c,a)+1),
$$
while
$$
\frac{∂}{∂ u(c,a)}\Bigl[∑_c^\primeη_u(c^\prime)\logη_u(c^\prime)\Bigr]=\logη_u(c)+1.
$$
Hence
$$
\frac{∂ J}{∂ u(c,a)}=\logη_u(c)-\log u(c,a).
$$ At the interior maximizer $u^⋆$ , the first-order condition on the fixed-support face reads
$$
0=\logη_u^⋆(c)-\log u^⋆(c,a)-∑_j=1^mλ_jG_j(c,a)-γ-ψ(c)+ψ(σ(c,a)).
$$
Rearranging and exponentiating gives the result on the active support. ∎
**Corollary 4.6 (Kernel representation on the active support)**
*On the active support where $η_u^⋆(c)>0$ , the induced kernel
$$
p^⋆(a\mid c):=\frac{u^⋆(c,a)}{η_u^⋆(c)}
$$
satisfies
$$
p^⋆(a\mid c)∝\exp\Bigl(-∑_j=1^mλ_jG_j(c,a)+ψ(σ(c,a))\Bigr), a∈ A.
$$
Equivalently,
$$
p^⋆(a\mid c)=\frac{\exp\Bigl(-∑_j=1^mλ_jG_j(c,a)+ψ(σ(c,a))\Bigr)}{∑_β∈ A\exp\Bigl(-∑_j=1^mλ_jG_j(c,β)+ψ(σ(c,β))\Bigr)}.
$$
If, in addition, the term $ψ(σ(c,a))$ is constant in $a$ for each fixed context $c$ , then this simplifies to the familiar rowwise exponential-family formula
$$
p^⋆(a\mid c)=\frac{\exp\bigl(-∑_j=1^mλ_jG_j(c,a)\bigr)}{∑_β∈ A\exp\bigl(-∑_j=1^mλ_jG_j(c,β)\bigr)}.
$$*
**Remark 4.7**
*In the general formula, the stationarity multipliers enter through $ψ(σ(c,a))$ , coupling rows across future contexts. The rowwise exponential-family form appears only in the special case of corollary 4.6.*
## 5 Local Geometry and the Gap Functional
Throughout this section, assume (L1)–(L2) and work on a fixed-support face. Let $(b,ξ)↦ u(b,ξ)$ denote a smooth local chart near $u^⋆=u(b_0,ξ_0)$ , where $b$ denotes retained observable coordinates and $ξ$ denotes fiber coordinates. Define
$$
\widetilde{J}(b,ξ):=J(u(b,ξ)). \tag{10}
$$
The selector map is then
$$
s(b)∈\arg\max_ξ\widetilde{J}(b,ξ). \tag{11}
$$
#### Intrinsic tangent space on a fixed face.
Fix an fixed-support face $\mathfrak{F}⊆Δ(A^r+1)$ and let $A_\mathfrak{F}$ denote the affine space of stationary-consistent block laws supported on $\mathfrak{F}$ and satisfying the retained observable constraints. Its translation space is
$$
Tan(A_\mathfrak{F})=≤ft\{h∈ℝ^A^{r+1}: \begin{aligned} &(h)⊆\mathfrak{F}, \textstyle∑_c,ah(c,a)=0,\\
&\textstyle∑_c,ah(c,a)G_j(c,a)=0 (j=1,\dots,m),\\
&\textstyle∑_a∈ Ah(c,a)=∑_a∈ Ah(a,c) (c∈ A^r)\end{aligned}\right\}. \tag{12}
$$
In particular, any local fiber coordinate $ξ$ on a fixed face may be chosen along a basis of this tangent space.
* Proof*
Inside a fixed face, the feasible set is cut out by the normalization constraint, the retained moment equalities, and the stationarity-consistency equations, all of which are affine-linear in the block-law coordinates. The translation space of that affine slice is therefore exactly the set of perturbations satisfying the corresponding homogeneous linear equations. The support restriction records that one remains on the same fixed-support face. ∎
#### Restricted Hessian in block-law coordinates.
Let $u$ be strictly positive on the fixed face $\mathfrak{F}$ , and let $h∈Tan(A_\mathfrak{F})$ . Then the second variation of the entropy-rate functional along $h$ is
$$
D^2J(u)[h,h]=-∑_c∈ A^r∑_a∈ A\frac{h(c,a)^2}{u(c,a)}+∑_c∈ A^r\frac{\bigl(∑_a∈ Ah(c,a)\bigr)^2}{η_u(c)}. \tag{13}
$$
Consequently $D^2J(u)[h,h]≤ 0$ for every such $h$ . If, in addition, $∑_a∈ Ah(c,a)=0$ for every context $c$ and $h≠ 0$ , then $D^2J(u)[h,h]<0$ .
* Proof*
Write
$$
J(u)=-∑_c,au(c,a)\log u(c,a)+∑_cη_u(c)\logη_u(c), η_u(c)=∑_au(c,a).
$$
For $u_t:=u+th$ , differentiation gives
$$
\frac{d^2}{dt^2}\Big|_t=0\Bigl[-∑_c,au_t(c,a)\log u_t(c,a)\Bigr]=-∑_c,a\frac{h(c,a)^2}{u(c,a)},
$$
while
$$
\frac{d^2}{dt^2}\Big|_t=0\Bigl[∑_cη_u_{t}(c)\logη_u_{t}(c)\Bigr]=∑_c\frac{\bigl(∑_ah(c,a)\bigr)^2}{η_u(c)}.
$$
Adding the two contributions yields the displayed formula. For each fixed context $c$ , Cauchy–Schwarz gives
$$
\Bigl(∑_ah(c,a)\Bigr)^2≤\Bigl(∑_au(c,a)\Bigr)\Bigl(∑_a\frac{h(c,a)^2}{u(c,a)}\Bigr)=η_u(c)∑_a\frac{h(c,a)^2}{u(c,a)}.
$$
Summing over $c$ proves that $D^2J(u)[h,h]≤ 0$ . If in addition each row sum $∑_ah(c,a)$ vanishes, then the positive term disappears and one obtains
$$
D^2J(u)[h,h]=-∑_c,a\frac{h(c,a)^2}{u(c,a)}<0
$$
for every nonzero $h$ . ∎
The fiber Hessian in theorems 5.4 and 5.7 is the restriction of $D^2J(u)$ to the tangent directions selected by the local coordinates. Concavity is automatic on each fixed face; strict negative definiteness is an additional hypothesis.
**Proposition 5.1 (Null directions of the restricted Hessian)**
*Let $u$ be strictly positive on the fixed face $\mathfrak{F}$ , and let $h∈Tan(A_\mathfrak{F})$ . Then the following are equivalent:
1. $D^2J(u)[h,h]=0$ .
1. For every context $c∈ A^r$ there exists a scalar $α(c)∈ℝ$ such that
$$
h(c,a)=α(c)u(c,a) for all a∈ A.
$$
In particular, the null directions of the restricted Hessian are exactly the row-rescaling directions that preserve each conditional law $p_u(·\mid c)$ .*
* Proof*
By the restricted-Hessian formula proved above,
$$
D^2J(u)[h,h]=-∑_c∈ A^r\Biggl[∑_a∈ A\frac{h(c,a)^2}{u(c,a)}-\frac{\bigl(∑_a∈ Ah(c,a)\bigr)^2}{η_u(c)}\Biggr].
$$
For each fixed context $c$ , set
$$
S_c:=∑_a∈ A\frac{h(c,a)^2}{u(c,a)}-\frac{\bigl(∑_a∈ Ah(c,a)\bigr)^2}{η_u(c)}.
$$
The Cauchy–Schwarz inequality applied to the vectors
$$
\Bigl(\frac{h(c,a)}{√{u(c,a)}}\Bigr)_a∈ A and \bigl(√{u(c,a)}\bigr)_a∈ A
$$
gives
$$
\Bigl(∑_a∈ Ah(c,a)\Bigr)^2≤η_u(c)∑_a∈ A\frac{h(c,a)^2}{u(c,a)},
$$
so each $S_c≥ 0$ . Therefore
$$
D^2J(u)[h,h]=-∑_c∈ A^rS_c≤ 0.
$$
Hence $D^2J(u)[h,h]=0$ if and only if $S_c=0$ for every context $c$ . Fix a context $c$ . Equality in Cauchy–Schwarz holds if and only if the two vectors displayed above are linearly dependent. Since $u(c,a)>0$ on the fixed face, this is equivalent to the existence of a scalar $α(c)$ such that
$$
\frac{h(c,a)}{√{u(c,a)}}=α(c)√{u(c,a)} for all a∈ A,
$$
which is exactly
$$
h(c,a)=α(c)u(c,a) for all a∈ A.
$$
Thus $S_c=0$ if and only if the row $h(c,·)$ is proportional to the row $u(c,·)$ . Since this must hold for every context, statements (i) and (ii) are equivalent. Finally, if $h(c,a)=α(c)u(c,a)$ for all $a$ , then for every sufficiently small $t$ such that $u+th$ remains in the face,
$$
p_u+th(a\mid c)=\frac{u(c,a)+th(c,a)}{η_u(c)+t∑_ah(c,a)}=\frac{(1+tα(c))u(c,a)}{(1+tα(c))η_u(c)}=p_u(a\mid c).
$$
Therefore these directions change only the row masses and preserve each conditional law. ∎
**Theorem 5.2 (Strict concavity on a fixed face)**
*Let $K$ be a convex subset of $A_\mathfrak{F}∩(\mathfrak{F})$ . Assume that for every $u∈ K$ and every nonzero direction $h∈Tan(A_\mathfrak{F})$ , there exists a context $c∈ A^r$ such that the row $h(c,·)$ is not proportional to the row $u(c,·)$ . Then $J$ is strictly concave on $K$ . In particular, if $K$ is nonempty and compact, then $J$ has a unique maximizer on $K$ .*
* Proof*
Let $u,v∈ K$ be distinct, and define
$$
h:=v-u.
$$
Because $K⊆A_\mathfrak{F}$ and $A_\mathfrak{F}$ is affine with translation space $Tan(A_\mathfrak{F})$ , one has
$$
h∈Tan(A_\mathfrak{F}) and h≠ 0.
$$
For each $t∈(0,1)$ , set
$$
ν_t:=(1-t)u+tv.
$$
Since $K$ is convex, $ν_t∈ K⊆A_\mathfrak{F}∩(\mathfrak{F})$ for every $t∈(0,1)$ . Consider the one-variable function
$$
φ(t):=J(ν_t).
$$
Because $J$ is $C^2$ on the relative interior of the fixed face, $φ$ is twice continuously differentiable on $(0,1)$ and
$$
φ^\prime\prime(t)=D^2J(ν_t)[h,h].
$$
By hypothesis, for each $t∈(0,1)$ the nonzero vector $h$ is not rowwise proportional to $ν_t$ in every context. Therefore proposition 5.1 yields
$$
D^2J(ν_t)[h,h]<0 for every t∈(0,1).
$$
Hence $φ$ is strictly concave on $(0,1)$ . In particular,
$$
φ(t)>(1-t)φ(0)+tφ(1) for every t∈(0,1), \tag{0}
$$
that is,
$$
J((1-t)u+tv)>(1-t)J(u)+tJ(v) for all distinct u,v∈ K, t∈(0,1).
$$
Thus $J$ is strictly concave on $K$ . If $K$ is nonempty and compact, continuity of $J$ on $K$ implies that $J$ attains its maximum on $K$ . Strict concavity on a convex set gives uniqueness of that maximizer. ∎
**Remark 5.3**
*theorem 5.2 is the fixed-face analogue of theorem 3.4: the global theorem is phrased in terms of pairs of feasible block laws, the fixed-face theorem in terms of the nullspace of the restricted Hessian.*
**Theorem 5.4 (Differentiability of the maximizer)**
*Assume $\widetilde{J}$ is $C^2$ near $(b_0,ξ_0)$ , where $ξ_0=s(b_0)$ , and that the chart is taken on a neighborhood with fixed active support. Assume also that $∂_ξ\widetilde{J}(b_0,ξ_0)=0$ and that the fiber Hessian
$$
∂_ξξ^2\widetilde{J}(b_0,ξ_0)
$$
is negative definite. Then there exists a neighborhood of $b_0$ in which the local selector is uniquely defined as the unique local maximizer on the fixed face, it is differentiable, and
$$
Ds(b_0)=-\bigl[∂_ξξ^2\widetilde{J}(b_0,ξ_0)\bigr]^-1∂_ξ b^2\widetilde{J}(b_0,ξ_0).
$$*
* Proof*
Define
$$
F(b,ξ):=∂_ξ\widetilde{J}(b,ξ).
$$
Because $\widetilde{J}$ is $C^2$ , the map $F$ is $C^1$ near $(b_0,ξ_0)$ . The selector is characterized in this chart by the first-order condition
$$
F(b,s(b))=0.
$$
The Jacobian of $F$ with respect to $ξ$ is
$$
∂_ξF(b_0,ξ_0)=∂_ξξ^2\widetilde{J}(b_0,ξ_0),
$$
which is nonsingular by hypothesis. Therefore the implicit function theorem yields a unique local $C^1$ critical-point map $s$ near $b_0$ . Because $∂_\ xi\ xi^2\ widetildeJ(b_0,\ xi_0)$ is negative definite, continuity of the Hessian implies that, after shrinking the neighborhood if necessary, the critical point $s(b)$ remains a strict local maximizer and is the unique local maximizer on the fixed face near $\ xi_0$ . Because the chart is taken on a neighborhood with fixed active support, no support change occurs inside this parametrization, so the theorem is purely local on one fixed face of the feasible set. Differentiating the identity $F(b,s(b))=0$ with respect to $b$ at $b_0$ gives
$$
∂_ξ b^2\widetilde{J}(b_0,ξ_0)+∂_ξξ^2\widetilde{J}(b_0,ξ_0)Ds(b_0)=0,
$$
and solving for $Ds(b_0)$ proves the formula. ∎
Define the optimized value function
$$
V(b):=\widetilde{J}(b,s(b)). \tag{14}
$$
**Proposition 5.5 (Envelope identities)**
*At the selected point,
$$
DV(b_0)=∂_b\widetilde{J}(b_0,ξ_0).
$$
Moreover,
$$
D^2V(b_0)=∂_bb^2\widetilde{J}(b_0,ξ_0)-∂_bξ^2\widetilde{J}(b_0,ξ_0)\bigl[∂_ξξ^2\widetilde{J}(b_0,ξ_0)\bigr]^-1∂_ξ b^2\widetilde{J}(b_0,ξ_0).
$$*
* Proof*
By the chain rule,
$$
DV(b)=∂_b\widetilde{J}(b,s(b))+∂_ξ\widetilde{J}(b,s(b))Ds(b).
$$
The second term vanishes by the selector first-order condition, which proves the first identity. Differentiating once more and substituting the formula for $Ds(b_0)$ yields the second identity. ∎
**Definition 5.6 (Gap functional)**
*Define
$$
Δ(b,ξ):=\widetilde{J}(b,s(b))-\widetilde{J}(b,ξ)≥ 0.
$$*
**Theorem 5.7 (Quadratic expansion of the gap)**
*Assume $\widetilde{J}$ is $C^2$ near $(b_0,ξ_0)$ , where $ξ_0=s(b_0)$ , and work in the fixed active-support chart introduced above. Let
$$
K(b_0):=-∂_ξξ^2\widetilde{J}(b_0,ξ_0),
$$
and assume that $K(b_0)$ is positive definite. Then
$$
Δ(b,ξ)=\frac{1}{2}(ξ-s(b))^⊤K(b_0)(ξ-s(b))+o(\|ξ-s(b)\|^2)
$$
as $(b,ξ)→(b_0,ξ_0)$ within this chart.*
* Proof*
Fix $b$ near $b_0$ and write $δ:=ξ-s(b)$ . Applying Taylor expansion to the map $ξ↦\widetilde{J}(b,ξ)$ around $ξ=s(b)$ gives
$$
\widetilde{J}(b,ξ)=\widetilde{J}(b,s(b))+∂_ξ\widetilde{J}(b,s(b))δ+\frac{1}{2}δ^⊤∂_ξξ^2\widetilde{J}(b,s(b))δ+o(\|δ\|^2).
$$
Because $s(b)$ is the local selector in this chart, the first-order term vanishes:
$$
∂_ξ\widetilde{J}(b,s(b))=0.
$$
Hence
$$
Δ(b,ξ)=-\frac{1}{2}δ^⊤∂_ξξ^2\widetilde{J}(b,s(b))δ+o(\|δ\|^2).
$$ Now add and subtract $K(b_0)$ :
$$
-∂_ξξ^2\widetilde{J}(b,s(b))=K(b_0)+\Bigl[-∂_ξξ^2\widetilde{J}(b,s(b))-K(b_0)\Bigr].
$$
Since $\widetilde{J}$ is $C^2$ and $s$ is continuous, the bracketed term tends to zero as $(b,ξ)→(b_0,ξ_0)$ within the fixed chart. Therefore
$$
δ^⊤\Bigl[-∂_ξξ^2\widetilde{J}(b,s(b))-K(b_0)\Bigr]δ=o(\|δ\|^2),
$$
which yields
$$
Δ(b,ξ)=\frac{1}{2}δ^⊤K(b_0)δ+o(\|δ\|^2).
$$
This is the claimed expansion. ∎
## 6 Hidden Realizations
Let $u^⋆$ be a selected stationary $(r+1)$ -block law with induced active-support kernel $p^⋆(a\mid c):=u^⋆(c,a)/η_u^⋆(c)$ .
**Theorem 6.1 (Random-mapping realization)**
*Let $u^⋆$ be a selected stationary $(r+1)$ -block law, and let $p^⋆(a\mid c)$ denote the induced kernel on the active contexts $(c∈ A^r with η_u^⋆(c)>0)$ . Then there exist a measurable map
$$
F:A^r×[0,1)→ A,
$$
an independent and identically distributed (i.i.d.) sequence $U_t∼uniform on [0,1)$ , and an initial context
$$
Y_-r^-1∼η_u^⋆,
$$
independent of $(U_t)_t≥ 0$ , such that the recursion
$$
Y_t=F(Y_t-r^t-1,U_t), t≥ 0,
$$
defines a stationary one-sided order- $r$ Markov chain whose stationary $(r+1)$ -block law is $u^⋆$ . In particular, it admits the two-sided stationary extension with the same block law.*
* Proof*
For each context $c∈ A^r$ , partition $[0,1)$ into half-open intervals
$$
\{I_c,a:a∈ A\}
$$
with lengths
$$
|I_c,a|=p^⋆(a\mid c).
$$
Define
$$
F(c,u)=a if u∈ I_c,a.
$$
Then, for every active context $c$ ,
$$
P(Y_t=a\mid Y_t-r^t-1=c)=P(U_t∈ I_c,a)=|I_c,a|=p^⋆(a\mid c).
$$
Thus the recursion has transition kernel $p^⋆$ . It remains to identify the stationary law. By definition,
$$
u^⋆(c,a)=η_u^⋆(c)p^⋆(a\mid c).
$$
Because $u^⋆$ is stationary-consistent, the context marginal $η_u^⋆$ is invariant for the induced context chain on $A^r$ . Choosing the initial context $Y_-r^-1$ with law $η_u^⋆$ therefore makes the resulting order- $r$ Markov chain stationary, and its stationary $(r+1)$ -block law is exactly $u^⋆$ . For contexts outside the active support, $F(c,·)$ may be chosen arbitrarily without affecting the realized stationary law. ∎
**Theorem 6.2 (Invariance under hidden measure-preserving actions)**
*Let $(Z_t)$ be a hidden state process, let $(U_t)$ be an independent and identically distributed (i.i.d.) sequence of $uniform on [0,1)$ variables independent of $(Z_t)$ , and let
$$
T_z:[0,1)→[0,1)
$$
be measurable and measure-preserving for each hidden state $z$ . Define
$$
ε_t:=T_Z_{t}(U_t), Y_t:=F(Y_t-r^t-1,ε_t).
$$
Let
$$
G_t:=σ(Z_s:s≤ t), F_t-1:=σ(Y_s:s≤ t-1).
$$
Then
$$
L(ε_t\midG_t\veeF_t-1)=uniform on [0,1).
$$
Consequently,
$$
L(ε_t\midG_t)=uniform on [0,1), L(ε_t\midF_t-1)=uniform on [0,1),
$$
and the induced visible kernel remains $p^⋆$ .*
* Proof*
For every Borel set $B⊆[0,1)$ , define
$$
A_B:=T_Z_{t}^-1(B).
$$
Then $A_B$ is $G_t$ -measurable, hence $G_t\veeF_t-1$ -measurable. Since $U_t$ is independent of $G_t\veeF_t-1$ and is uniform on $[0,1)$ , the conditional-expectation identity for indicators of random measurable sets yields
$$
P\bigl(U_t∈ A_B\midG_t\veeF_t-1\bigr)=λ(A_B) a.s.
$$
Because $ε_t=T_Z_{t}(U_t)$ , this becomes
$$
P(ε_t∈ B\midG_t\veeF_t-1)=λ\bigl(T_Z_{t}^-1(B)\bigr).
$$
Since each $T_z$ preserves Lebesgue measure,
$$
λ\bigl(T_Z_{t}^-1(B)\bigr)=λ(B).
$$
Therefore
$$
L(ε_t\midG_t\veeF_t-1)=uniform on [0,1).
$$
The two displayed marginal conditional-law statements follow immediately by the tower property. Finally, for every context $c∈ A^r$ and symbol $a∈ A$ ,
$$
P(Y_t=a\mid Y_t-r^t-1=c)=P(ε_t∈ I_c,a\mid Y_t-r^t-1=c)=|I_c,a|=p^⋆(a\mid c).
$$
Thus hidden structural action is compatible with no change in the visible law while leaving the visible transition kernel unchanged. ∎
**Remark 6.3**
*The theorem shows that hidden structure may act on the primitive shocks in a measure-preserving way while remaining invisible at the level of the conditional visible law. In particular, this section proves a realizability and invariance statement for the selected visible kernel, not a realization theorem inside a prescribed hidden observational fiber.*
## 7 Empirical Maximizers and Consistency
Let
$$
Y_0,\dots,Y_n
$$
be one observed visible sample path on the finite alphabet $A$ . For $c∈ A^r$ and $a∈ A$ , define the empirical $(r+1)$ -block frequency
$$
\hat{u}_n(c,a):=\frac{1}{n-r}∑_t=r^n-11\{Y_t-r^t-1=c, Y_t=a\}. \tag{15}
$$
Its empirical context marginal is
$$
\hat{η}_n(c):=∑_a∈ A\hat{u}_n(c,a).
$$
If the experiment-induced features are $G_1,\dots,G_m$ , define the empirical targets by
$$
\hat{b}_n,j:=∑_c,a\hat{u}_n(c,a)G_j(c,a), j=1,\dots,m.
$$
For a deterministic target vector $b∈ℝ^m$ , write
$$
U(b):=\Bigl\{u on A^r+1:u is stationary-consistent and ∑_c,au(c,a)G_j(c,a)=b_j, j=1,\dots,m\Bigr\}.
$$
Then the empirical feasible class is
$$
\hat{U}_n=U(\hat{b}_n).
$$
Define the empirical criterion
$$
\hat{J}_n(u):=-∑_c,au(c,a)\log\frac{u(c,a)}{η_u(c)}.
$$
Because the objective depends only on $u$ and not explicitly on $n$ , one may equivalently write
$$
\hat{J}_n(u)=J(u) on \hat{U}_n.
$$
Let
$$
\hat{u}_n^⋆∈\arg\max_u∈\hat{U_n}J(u) \tag{16}
$$
be any empirical selector.
**Proposition 7.1 (Empirical block-frequency convergence)**
*Assume the true visible process is stationary, irreducible, and $r$ -step Markov on the finite alphabet $A$ . Then for each $(c,a)∈ A^r+1$ ,
$$
\hat{u}_n(c,a)→ u_0(c,a) almost surely,
$$
where $u_0$ denotes the true stationary $(r+1)$ -block law. Consequently,
$$
\hat{η}_n(c)→η_u_0(c) almost surely
$$
for every context $c∈ A^r$ .*
* Proof*
For each fixed block $(c,a)$ , the indicator
$$
1\{Y_t-r^t-1=c, Y_t=a\}
$$
is bounded. The ergodic theorem for finite-state Markov chains therefore gives
$$
\hat{u}_n(c,a)→ u_0(c,a) almost surely
$$
for each $(c,a)$ . Summing over $a∈ A$ yields the convergence of the empirical context marginals. ∎
**Proposition 7.2 (Empirical feature convergence)**
*Assume the hypotheses of proposition 7.1. Then for each $j=1,\dots,m$ ,
$$
\hat{b}_n,j→ b_0,j:=∑_c,au_0(c,a)G_j(c,a) almost surely.
$$
Equivalently,
$$
\hat{b}_n→ b_0 almost surely.
$$*
* Proof*
Because $A^r+1$ is finite, each feature $G_j$ is bounded on $A^r+1$ . Hence
$$
\hat{b}_n,j-b_0,j=∑_c,a\bigl(\hat{u}_n(c,a)-u_0(c,a)\bigr)G_j(c,a),
$$
and every summand converges almost surely to zero by proposition 7.1. Therefore $\hat{b}_n,j→ b_0,j$ almost surely for each $j$ , and thus $\hat{b}_n→ b_0$ almost surely. ∎
Fix the fixed-support face
$$
\mathfrak{F}⊆Δ(A^r+1)
$$
containing the selected point $u^⋆$ in its relative interior, and let
$$
A_\mathfrak{F}
$$
denote the affine subspace of stationary-consistent block laws supported on $\mathfrak{F}$ . Write
$$
Tan(A_\mathfrak{F})
$$
for the translation space of the affine set $A_\mathfrak{F}$ . Define the linear moment map
$$
T:A_\mathfrak{F}→ℝ^m, T(u):=\Bigl(∑_c,au(c,a)G_j(c,a)\Bigr)_j=1^m.
$$
**Lemma 7.3 (Local feasible continuation near the selected point)**
*Assume $u^⋆∈U(b_0)∩A_\mathfrak{F}$ lies in the relative interior of the face $\mathfrak{F}$ , and assume that the restriction of $T$ to the translation space of $A_\mathfrak{F}$ has full row rank. Then there exist a neighborhood $B$ of $b_0$ in $ℝ^m$ , a neighborhood $N$ of $u^⋆$ in $A_\mathfrak{F}$ , and a constant $L>0$ such that
$$
\overline{N}⊂(\mathfrak{F}),
$$
where $$ denotes relative interior, and for every $b∈ B$ there exists a point
$$
u(b)∈U(b)∩ N
$$
satisfying
$$
\|u(b)-u^⋆\|≤ L\|b-b_0\|.
$$
In particular,
$$
U(b)∩ N≠∅
$$
for every $b∈ B$ .*
* Proof*
Because the restricted moment map has full row rank, there exists a linear right inverse
$$
R:ℝ^m→Tan(A_\mathfrak{F})
$$
with $TR=I_m$ . Define
$$
u(b):=u^⋆+R(b-b_0).
$$
Then $u(b)∈A_\mathfrak{F}$ and
$$
T\bigl(u(b)\bigr)=T(u^⋆)+TR(b-b_0)=b.
$$
Thus $u(b)∈U(b)∩A_\mathfrak{F}$ whenever it remains in the face. Since $u^⋆∈(\mathfrak{F})$ , there exists a neighborhood $N$ of $u^⋆$ in $A_\mathfrak{F}$ whose compact closure satisfies
$$
\overline{N}⊂(\mathfrak{F}).
$$
Because $u(b)→ u^⋆$ as $b→ b_0$ , there exists a neighborhood $B$ of $b_0$ such that $u(b)∈ N$ for every $b∈ B$ . This gives the claim, with $L:=\|R\|$ . ∎
**Theorem 7.4 (Local consistency on a fixed-support face)**
*Assume the hypotheses of propositions 7.1 and 7.2. Assume also that:
1. the population selector $u^⋆$ is the unique maximizer of $J$ over
$$
U(b_0)∩A_\mathfrak{F},
$$
1. $u^⋆$ lies in the relative interior of the fixed-support face $\mathfrak{F}$ ,
1. the restricted moment map in lemma 7.3 has full row rank.
Then there exists a neighborhood $N$ of $u^⋆$ in $A_\mathfrak{F}$ with compact closure contained in $(\mathfrak{F})$ such that the following holds almost surely for all sufficiently large $n$ : the compact local empirical feasible class $\hat{U}_n∩\overline{N}$ is nonempty, and every empirical maximizer
$$
\hat{u}_n^⋆∈\arg\max_u∈\hat{U_n∩\overline{N}}J(u)
$$
satisfies $\hat{u}_n^⋆→ u^⋆$ almost surely.*
* Proof*
Because $u^⋆$ is the unique maximizer of the continuous function $J$ on the compact set
$$
U(b_0)∩A_\mathfrak{F},
$$
there exists a neighborhood $N$ of $u^⋆$ in $A_\mathfrak{F}$ with compact closure
$$
\overline{N}⊂(\mathfrak{F})
$$
and an $ε>0$ such that
$$
J(u)≤ J(u^⋆)-ε for every u∈\bigl(U(b_0)∩A_\mathfrak{F}\bigr)∖ N.
$$
Shrinking $N$ if necessary, I may also assume that lemma 7.3 holds on this same neighborhood. By proposition 7.2, one has [5, 4]
$$
\hat{b}_n→ b_0 almost surely.
$$
Hence, by lemma 7.3, for all sufficiently large $n$ there exists
$$
u_n∈\hat{U}_n∩ N
$$
with $u_n→ u^⋆$ . In particular,
$$
\hat{U}_n∩ N≠∅
$$
eventually almost surely. Fix an almost sure sample point on which $\hat{b}_n→ b_0$ . Let $\hat{u}_n^⋆$ be any local empirical maximizer in $\hat{U}_n∩ N$ . Because $\overline{N}$ is compact, every subsequence of $(\hat{u}_n^⋆)$ has a further convergent subsequence, let
$$
\hat{u}_n_{k}^⋆→\bar{u}∈\overline{N}.
$$
Since
$$
T(\hat{u}_n_{k}^⋆)=\hat{b}_n_{k} and \hat{b}_n_{k}→ b_0,
$$
continuity of $T$ yields
$$
T(\bar{u})=b_0,
$$
so $\bar{u}∈U(b_0)∩\overline{N}$ . Moreover, because $u_n_{k}∈\hat{U}_n_{k}∩ N$ and $\hat{u}_n_{k}^⋆$ maximizes $J$ over that local feasible class,
$$
J(\hat{u}_n_{k}^⋆)≥ J(u_n_{k}).
$$
Passing to the limit and using continuity of $J$ gives
$$
J(\bar{u})≥ J(u^⋆).
$$
Since $u^⋆$ is feasible for the population problem and is the unique maximizer of $J$ on $U(b_0)∩A_\mathfrak{F}$ , it follows that $\bar{u}=u^⋆$ . Thus every convergent subsequence of $(\hat{u}_n^⋆)$ converges to $u^⋆$ , and therefore the whole sequence converges almost surely to $u^⋆$ . ∎
The local asymptotic statistic from theorem 5.7 may be evaluated at empirical local coordinates derived from $\hat{u}_n^⋆$ . Thus the gap functional is a population-geometric object and may serve as the basis of a testing procedure under additional asymptotic assumptions.
**Remark 7.5**
*A full empirical theory across support changes would require substantially more work. The present result isolates the local argument: convergence of block frequencies yields convergence of retained moments, full row rank yields nearby feasible continuation on one face, and uniqueness of the population selector yields the separation needed to force local empirical maximizers back to $u^⋆$ .*
## 8 An Aliased Hidden-State Example
This section gives a concrete application in which the selector resolves a genuinely underidentified visible model, while hidden completion remains non-unique. The construction stays within the finite-state finite-memory scope of the paper and makes the observational-fiber interpretation explicit.
Let the hidden state space be
$$
E:=\{a_0,a_1,b_0,b_1\},
$$
and let the visible alphabet be $A:=\{0,1\}$ . Define the observation map
$$
φ(a_0)=φ(a_1)=0, φ(b_0)=φ(b_1)=1.
$$
Thus two hidden states are observationally aliased into each visible symbol.
Fix parameters
$$
a,b,λ,μ∈(0,1)
$$
and define a Markov chain $(X_t)$ on $E$ by the transition matrix
$$
K=\begin{pmatrix}(1-a)λ&(1-a)(1-λ)&aμ&a(1-μ)\\
(1-a)λ&(1-a)(1-λ)&aμ&a(1-μ)\\
bλ&b(1-λ)&(1-b)μ&(1-b)(1-μ)\\
bλ&b(1-λ)&(1-b)μ&(1-b)(1-μ)\end{pmatrix},
$$
where the rows are ordered as $(a_0,a_1,b_0,b_1)$ . Let the visible process be
$$
Y_t:=φ(X_t).
$$
The point of the construction is that the visible transition probabilities depend only on the aggregated parameters $a$ and $b$ , whereas the hidden completion still depends on the latent splitting parameters $λ$ and $μ$ .
**Proposition 8.1 (Aliased hidden realization of any binary first-order chain)**
*Let
$$
P=\begin{pmatrix}1-a&a\\
b&1-b\end{pmatrix} with a,b∈(0,1)
$$
be a stationary irreducible binary Markov kernel, and let
$$
m:=\frac{a}{a+b}.
$$
For any $λ,μ∈(0,1)$ , the hidden transition matrix $K$ defined above has the following properties:
1. The hidden chain is irreducible and aperiodic.
1. It admits the stationary distribution
$$
π^E=\bigl((1-m)λ,(1-m)(1-λ),mμ,m(1-μ)\bigr).
$$
1. The observed process $Y_t=φ(X_t)$ is a stationary binary first-order Markov chain with transition matrix $P$ .
In particular, every stationary irreducible binary first-order Markov chain admits infinitely many aliased hidden-state realizations of this form.*
* Proof*
Because all entries of $K$ are strictly positive, the hidden chain is irreducible and aperiodic. This proves (i). To prove (ii), write
$$
π^E=(π_a_0,π_a_1,π_b_0,π_b_1):=\bigl((1-m)λ,(1-m)(1-λ),mμ,m(1-μ)\bigr).
$$
It is a probability vector because
$$
(1-m)λ+(1-m)(1-λ)+mμ+m(1-μ)=1.
$$
Now compute the first coordinate of $π^EK$ :
| | $\displaystyle(π^EK)_a_0$ | $\displaystyle=\bigl(π_a_0+π_a_1\bigr)(1-a)λ+\bigl(π_b_0+π_b_1\bigr)bλ$ | |
| --- | --- | --- | --- |
Because $m=a/(a+b)$ , one has
$$
(1-m)a=mb,
$$
so
$$
(1-m)(1-a)+mb=(1-m)-(1-m)a+mb=1-m.
$$
Hence
$$
(π^EK)_a_0=(1-m)λ=π_a_0.
$$
The same calculation gives
$$
(π^EK)_a_1=(1-m)(1-λ)=π_a_1.
$$
Likewise,
| | $\displaystyle(π^EK)_b_0$ | $\displaystyle=\bigl(π_a_0+π_a_1\bigr)aμ+\bigl(π_b_0+π_b_1\bigr)(1-b)μ$ | |
| --- | --- | --- | --- |
where the identity $(1-m)a=mb$ was used again. The fourth coordinate is analogous, so
$$
π^EK=π^E.
$$
Thus $π^E$ is stationary. To prove (iii), note that the visible transition probabilities are obtained by summing the hidden transition probabilities over each fiber of $φ$ . If the current hidden state is $a_0$ , then the probability of moving to the visible symbol $0$ is
$$
K(a_0,a_0)+K(a_0,a_1)=(1-a)λ+(1-a)(1-λ)=1-a,
$$
and the probability of moving to the visible symbol $1$ is
$$
K(a_0,b_0)+K(a_0,b_1)=aμ+a(1-μ)=a.
$$
Exactly the same identities hold when the current hidden state is $a_1$ , because the first two rows of $K$ are equal. Similarly, if the current hidden state is $b_0$ or $b_1$ , then the last two rows of $K$ give
$$
P(Y_t+1=0\mid X_t=x)=b, P(Y_t+1=1\mid X_t=x)=1-b, x∈\{b_0,b_1\}.
$$
Consequently, for each visible state $y∈\{0,1\}$ and each symbol $z∈\{0,1\}$ , the quantity
$$
P(Y_t+1=z\mid X_t=x)
$$
is the same for all hidden states $x$ satisfying $φ(x)=y$ , and its common value is exactly $P(y,z)$ . By conditioning on the sigma-field generated by $Y_t$ and applying the tower property, it follows that
$$
P(Y_t+1=z\mid Y_t=y)=P(y,z).
$$
Hence $(Y_t)$ is a binary first-order Markov chain with transition matrix $P$ . Since the hidden chain is stationary under $π^E$ , the visible process is stationary as well. This proves (iii). Finally, varying $λ$ and $μ$ in $(0,1)$ changes the hidden transition matrix $K$ while leaving the visible transition matrix $P$ unchanged. Hence one obtains infinitely many distinct aliased hidden-state realizations of the same visible chain. ∎
**Corollary 8.2 (Visible underidentification under a fixed mean)**
*Fix $m∈(0,1)$ . For each
$$
q∈\bigl(0,\min\{m,1-m\}\bigr),
$$
define
$$
a(q):=\frac{q}{1-m}, b(q):=\frac{q}{m}.
$$
Then for every choice of $λ,μ∈(0,1)$ , the aliased hidden transition matrix obtained from
$$
a=a(q), b=b(q)
$$
induces a stationary binary visible first-order Markov chain with stationary mean $m$ . As $q$ varies, these visible laws are distinct. Consequently, if the retained visible observable is only the stationary mean $m$ , then the visible completion problem is non-unique even inside this single aliased hidden-state architecture.*
* Proof*
For each admissible $q$ , the definitions of $a(q)$ and $b(q)$ give
$$
(1-m)a(q)=q=mb(q),
$$
so the stationary mean of the visible chain in proposition 8.1 is exactly $m$ . Distinct values of $q$ give distinct transition matrices
$$
\begin{pmatrix}1-a(q)&a(q)\\
b(q)&1-b(q)\end{pmatrix},
$$
hence distinct visible laws. Therefore the visible law is not determined by the retained mean alone. ∎
**Theorem 8.3 (Visible maximizing completion under aliasing)**
*Fix $m∈(0,1)$ , and let $V_m^aliased$ denote the class of stationary irreducible binary visible first-order Markov laws generated by the aliased hidden-state construction of propositions 8.1 and 8.2 under the single retained visible observable
$$
P(Y_t=1)=m.
$$
Then the entropy-rate functional has a unique maximizer on $V_m^aliased$ . This unique selector is the i.i.d. Bernoulli $(m)$ law, equivalently the binary first-order Markov chain with transition matrix
$$
P^⋆=\begin{pmatrix}1-m&m\\
1-m&m\end{pmatrix}.
$$
Thus the selector canonically resolves the visible underidentification inside this aliased hidden-state architecture, even though the hidden completion remains non-unique in the explicit sense of proposition 8.4, equivalently, the selected visible law still has an infinite hidden observational fiber by corollary 8.6.*
* Proof*
By corollary 8.2, the class $V_m^aliased$ is parameterized by
$$
q∈\bigl(0,\min\{m,1-m\}\bigr)
$$
through
$$
a(q)=\frac{q}{1-m}, b(q)=\frac{q}{m}.
$$
For such a chain the entropy rate is
$$
h(q)=(1-m)h_2\Bigl(\frac{q}{1-m}\Bigr)+mh_2\Bigl(\frac{q}{m}\Bigr),
$$
where
$$
h_2(p):=-p\log p-(1-p)\log(1-p).
$$
Since
$$
h_2^\prime(p)=\log\frac{1-p}{p} and h_2^\prime\prime(p)=-\frac{1}{p(1-p)}<0 (0<p<1),
$$
one obtains
| | $\displaystyle h^\prime(q)$ | $\displaystyle=h_2^\prime\Bigl(\frac{q}{1-m}\Bigr)+h_2^\prime\Bigl(\frac{q}{m}\Bigr)$ | |
| --- | --- | --- | --- |
and therefore
$$
h^\prime\prime(q)=-\frac{1}{1-m-q}-\frac{1}{m-q}-\frac{2}{q}<0 for q∈\bigl(0,\min\{m,1-m\}\bigr).
$$
Hence $h$ is strictly concave on the whole parameter interval, so it has at most one maximizer. The critical-point equation $h^\prime(q)=0$ is equivalent to
$$
(1-m-q)(m-q)=q^2.
$$
Expanding the left-hand side gives
$$
m(1-m)-q+q^2=q^2,
$$
so the unique critical point is
$$
q^⋆=m(1-m).
$$
Because
$$
0<m(1-m)<\min\{m,1-m\} for every m∈(0,1),
$$
this critical point lies in the admissible interval. Strict concavity now implies that $q^⋆$ is the unique global maximizer of $h$ on $V_m^aliased$ . Substituting $q^⋆$ into the parameterization yields
$$
a^⋆=\frac{q^⋆}{1-m}=m, b^⋆=\frac{q^⋆}{m}=1-m.
$$
Therefore the unique maximizing visible kernel is
$$
P^⋆=\begin{pmatrix}1-m&m\\
1-m&m\end{pmatrix},
$$
which is the i.i.d. Bernoulli $(m)$ law. This proves the visible-selection statement. The hidden non-uniqueness assertion is then made precise by propositions 8.4 and 8.6. ∎
**Proposition 8.4 (Failure of hidden maximizing completion)**
*Fix $m∈(0,1)$ and let $P^⋆$ be the selected visible kernel from theorem 8.3. Then the class of aliased hidden-state realizations that generate $P^⋆$ contains infinitely many distinct hidden transition matrices, namely
$$
K_λ,μ=\begin{pmatrix}(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\end{pmatrix}, λ,μ∈(0,1).
$$
Consequently, the selector of the present paper determines a canonical visible completion, but it does not determine a canonical hidden completion.*
* Proof*
Set $a=m$ and $b=1-m$ in proposition 8.1. Then the resulting visible kernel is exactly
$$
P^⋆=\begin{pmatrix}1-m&m\\
1-m&m\end{pmatrix}.
$$
For every pair $(λ,μ)∈(0,1)^2$ , proposition 8.1 shows that the hidden transition matrix $K_λ,μ$ realizes this same visible law. It remains to show that different parameter pairs produce different hidden completions. If
$$
(λ,μ)≠(λ^\prime,μ^\prime),
$$
then either $λ≠λ^\prime$ or $μ≠μ^\prime$ . In the first case the entry in position $(a_0,a_0)$ differs:
$$
K_λ,μ(a_0,a_0)=(1-m)λ≠(1-m)λ^\prime=K_λ^\prime,μ^{\prime}(a_0,a_0).
$$
In the second case the entry in position $(a_0,b_0)$ differs:
$$
K_λ,μ(a_0,b_0)=mμ≠ mμ^\prime=K_λ^\prime,μ^{\prime}(a_0,b_0).
$$
Hence
$$
K_λ,μ≠ K_λ^\prime,μ^{\prime}.
$$
Since $(0,1)^2$ contains infinitely many points, this yields infinitely many distinct hidden transition matrices realizing the same selected visible law. Therefore the entropy selector canonically resolves the visible ambiguity but leaves the hidden completion problem non-unique. This is exactly the distinction between visible selection and hidden completion that motivates the observational-fiber viewpoint of the paper. ∎
**Proposition 8.5 (Uniqueness of the visible maximizer and nonuniqueness of hidden realizations)**
*Fix $m∈(0,1)$ . Within the aliased hidden-state architecture of this section, the retained visible observable
$$
P(Y_t=1)=m
$$
determines a unique canonical visible completion by theorem 8.3, but it does not determine a unique hidden completion by proposition 8.4. Consequently, any rule that selects a single hidden law from this architecture must use additional information not contained in the retained visible observable and the observation map alone.*
* Proof*
The first assertion is exactly the content of theorem 8.3. The second assertion is exactly the content of proposition 8.4. Therefore the observation experiment together with the retained visible mean determines a unique selected visible law but leaves infinitely many admissible hidden realizations. Suppose, to the contrary, that there were a hidden-selection rule on this architecture determined solely by the observation map and the retained visible observable. Applied at the present value of $m$ , such a rule would have to choose one hidden law from a family of infinitely many hidden laws that are indistinguishable at the retained visible level. Hence the rule would require a tie-breaking criterion not encoded in the given visible information. This contradicts the assumption that the rule is determined by the observation map and retained visible observable alone. ∎
**Corollary 8.6 (Continuum hidden observational fiber at the selected visible law)**
*Let $ν^⋆$ denote the stationary visible path law of the selected Bernoulli $(m)$ process from theorem 8.3, and let $Π$ be the pathwise observation map induced by $φ$ . For each $(λ,μ)∈(0,1)^2$ , let $Q_λ,μ$ be the stationary path law of the hidden Markov chain with transition matrix $K_λ,μ$ from proposition 8.4. Then
$$
Q_λ,μ∈E_Π(ν^⋆) for every (λ,μ)∈(0,1)^2,
$$
and the set
$$
\{Q_λ,μ:(λ,μ)∈(0,1)^2\}
$$
is infinite. In particular, the observational fiber of the selected visible law is non-singleton and in fact contains infinitely many stationary hidden laws.*
* Proof*
By proposition 8.4, every matrix $K_λ,μ$ generates the same selected visible kernel $P^⋆$ . Because each such hidden chain is finite-state, irreducible, and aperiodic, it has a unique stationary path law, denoted here by $Q_λ,μ$ . The visible process obtained from $Q_λ,μ$ under the pathwise observation map $Π$ is exactly the stationary Bernoulli $(m)$ law $ν^⋆$ . Hence
$$
Π_\#Q_λ,μ=ν^⋆,
$$
so by the definition of the observational fiber in section 2, one has
$$
Q_λ,μ∈E_Π(ν^⋆).
$$
If $(λ,μ)≠(λ^\prime,μ^\prime)$ , then proposition 8.4 gives
$$
K_λ,μ≠ K_λ^\prime,μ^{\prime}.
$$
Because every state has strictly positive stationary mass, the stationary path law of a finite-state Markov chain determines its two-point cylinder probabilities and hence its transition matrix. Therefore distinct transition matrices yield distinct stationary path laws, so
$$
Q_λ,μ≠ Q_λ^\prime,μ^{\prime}.
$$
Therefore the family above is infinite. ∎
### 8.1 Comparison with Related Approaches
The selector is related to three nearby objectives. It concerns canonical visible completion rather than hidden identification. It is not obtained by projection alone from a hidden chain, but by entropy-rate maximization on a visible feasible class determined by retained observables. The exponential representation appears, when it appears, as a consequence of the constrained block-law optimization itself.
**Proposition 8.7 (Comparison with hidden identification)**
*Fix $m∈(0,1)$ and work in the aliased hidden-state architecture of proposition 8.1. Then the following statements hold.
1. The retained visible observable $P(Y_t=1)=m$ does not identify a unique hidden law.
1. The selected visible law is not determined by projection alone from the retained visible observable: within the same aliased architecture and under the same retained mean $m$ , there exist infinitely many distinct visible first-order Markov laws, but exactly one entropy-rate maximizer.
Consequently, the selector of the present paper solves a canonical visible-completion problem and not a hidden-identification problem.*
* Proof*
Assertion (i) is exactly the content of corollary 8.6: the selected visible law $ν^⋆$ has an infinite observational fiber
$$
E_Π(ν^⋆).
$$
In particular, the retained visible observable together with the observation map does not determine a unique hidden stationary law. For (ii), corollary 8.2 shows that under the fixed retained mean $m$ there are infinitely many distinct visible first-order Markov laws in the same aliased architecture, parameterized by
$$
q∈\bigl(0,\min\{m,1-m\}\bigr).
$$
By theorem 8.3, the entropy-rate functional has a unique maximizer on this visible class, attained at
$$
q^⋆=m(1-m),
$$
which yields the Bernoulli $(m)$ law. Therefore the selected visible law is not obtained from the retained mean by projection alone, because the same retained mean is compatible with infinitely many projected visible laws. What singles out one visible law is the additional canonical entropy-rate selection principle. Combining (i) and (ii) proves the final statement. ∎
**Proposition 8.8 (Comparison with exponential-family selection)**
*Assume the hypotheses of proposition 4.5. Then any exponential representation of the selected active-support kernel in the present paper is derived from the constrained visible optimization problem itself. More precisely, on the active support one has the multiplier identity
$$
\log\frac{u^⋆(c,a)}{η_u^⋆(c)}=-γ-∑_j=1^mλ_jG_j(c,a)-ψ(c)+ψ(σ(c,a)),
$$
so the selected kernel is determined by the block-law KKT system. In particular, the paper does not begin by specifying an external exponential-family transition model on the hidden side.*
* Proof*
The displayed identity is exactly the conclusion of proposition 4.5. Therefore any exponential expression for the selected visible kernel is obtained from the Lagrange multipliers attached to the normalization, stationarity, and retained-observable constraints of the visible block-law problem. It follows that the exponential structure is derived rather than postulated. In particular, in the aliased hidden-state application of this paper, propositions 8.4 and 8.6 show that even the selected visible law is compatible with infinitely many distinct hidden laws. Hence an externally imposed hidden exponential-family ansatz would add modeling structure not determined by the retained visible observable and observation map alone. ∎
**Proposition 8.9 (Comparison with hidden-entropy maximization)**
*Fix $m∈(0,1)$ and consider the family of hidden completions
$$
K_λ,μ=\begin{pmatrix}(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\\
(1-m)λ&(1-m)(1-λ)&mμ&m(1-μ)\end{pmatrix}, (λ,μ)∈(0,1)^2,
$$
which all generate the selected visible Bernoulli $(m)$ law from theorem 8.3. Then the hidden entropy rate of the stationary hidden chain equals
$$
H_hid(λ,μ)=h_2(m)+(1-m)h_2(λ)+mh_2(μ).
$$
In particular, hidden entropy maximization over this completion family has the unique maximizer
$$
λ^⋆=μ^⋆=\frac{1}{2}.
$$
Therefore hidden entropy maximization defines a different selection problem from the visible entropy-rate selector of the present paper: it chooses one hidden completion only after introducing an additional hidden-level objective.*
* Proof*
For the selected visible law one has $a=m$ and $b=1-m$ , so every row of $K_λ,μ$ is the same probability vector
$$
r_λ,μ:=\bigl((1-m)λ,(1-m)(1-λ),mμ,m(1-μ)\bigr).
$$
Hence the hidden chain is i.i.d. with common one-step law $r_λ,μ$ . In particular, its entropy rate equals the Shannon entropy of this single-step distribution:
$$
H_hid(λ,μ)=-∑_x∈ Er_λ,μ(x)\log r_λ,μ(x).
$$
Expanding the four terms gives
| | $\displaystyle H_hid(λ,μ)$ | $\displaystyle=-(1-m)λ\log\bigl((1-m)λ\bigr)-(1-m)(1-λ)\log\bigl((1-m)(1-λ)\bigr)$ | |
| --- | --- | --- | --- |
Grouping the $\log(1-m)$ , $\log m$ , $\logλ$ , $\log(1-λ)$ , $\logμ$ , and $\log(1-μ)$ terms yields
$$
H_hid(λ,μ)=h_2(m)+(1-m)h_2(λ)+mh_2(μ).
$$
Since the binary entropy function is strictly concave on $(0,1)$ and has its unique maximizer at $1/2$ , the weighted sum above is strictly concave in $(λ,μ)$ and is uniquely maximized at
$$
λ^⋆=μ^⋆=\frac{1}{2}.
$$
This maximizer depends on the hidden splitting coordinates $λ$ and $μ$ , which are invisible at the retained visible level. Therefore maximizing hidden entropy over hidden completions is a different problem from selecting the canonical visible completion by visible entropy-rate maximization. ∎
**Proposition 8.10 (Reduction to visible optimization)**
*Let
$$
S:P_stat(H^ℤ)→ℝ
$$
be an observable-determined real-valued criterion in the sense of proposition 8.18. Let $\widetilde{S}$ be the unique visible-law functional satisfying
$$
S=\widetilde{S}∘Π_\#.
$$
Then, for every nonempty class
$$
C⊆P_stat(H^ℤ),
$$
one has
$$
\sup_Q∈CS(Q)=\sup_ν∈Π_{\#C}\widetilde{S}(ν).
$$
Moreover, a hidden law $Q^⋆∈C$ maximizes $S$ over $C$ if and only if its visible law $ν^⋆:=Π_\#Q^⋆$ maximizes $\widetilde{S}$ over $Π_\#C$ . In particular, observable-determined optimization can determine at most a visible law, it cannot distinguish two hidden laws in $C$ that lie in the same observational fiber.*
* Proof*
By proposition 8.18, there exists a unique map
$$
\widetilde{S}:Π_\#\bigl(P_stat(H^ℤ)\bigr)→ℝ
$$
such that
$$
S(Q)=\widetilde{S}(Π_\#Q) for every Q∈P_stat(H^ℤ).
$$
Therefore
$$
\{S(Q):Q∈C\}=\{\widetilde{S}(ν):ν∈Π_\#C\}.
$$
Taking suprema gives
$$
\sup_Q∈CS(Q)=\sup_ν∈Π_{\#C}\widetilde{S}(ν).
$$
This proves the first assertion. Now let $Q^⋆∈C$ , and write $ν^⋆:=Π_\#Q^⋆$ . If $Q^⋆$ maximizes $S$ over $C$ , then for every $ν∈Π_\#C$ there exists some $Q∈C$ with $Π_\#Q=ν$ . Hence
$$
\widetilde{S}(ν)=S(Q)≤ S(Q^⋆)=\widetilde{S}(ν^⋆),
$$
so $ν^⋆$ maximizes $\widetilde{S}$ over $Π_\#C$ . Conversely, if $ν^⋆$ maximizes $\widetilde{S}$ over $Π_\#C$ , then for every $Q∈C$ ,
$$
S(Q)=\widetilde{S}(Π_\#Q)≤\widetilde{S}(ν^⋆)=S(Q^⋆),
$$
so $Q^⋆$ maximizes $S$ over $C$ . This proves the equivalence. Finally, if $Q_1,Q_2∈C$ satisfy
$$
Π_\#Q_1=Π_\#Q_2,
$$
then
$$
S(Q_1)=\widetilde{S}(Π_\#Q_1)=\widetilde{S}(Π_\#Q_2)=S(Q_2).
$$
Thus observable-determined optimization cannot distinguish points in the same observational fiber. ∎
**Corollary 8.11 (Maximality of visible selection)**
*Within the framework of this paper, any observable-determined optimization principle can canonically determine at most a visible law. In particular, on the selected Bernoulli fiber of theorem 8.3, the canonical visible selector is the finest resolution obtainable without adding hidden-side information beyond the observation experiment and retained visible observables.*
* Proof*
The first statement is exactly the final assertion of proposition 8.10. The second follows by applying that proposition to the non-singleton selected fiber identified in corollary 8.6: because the fiber contains infinitely many hidden laws with the same visible image, no observable-determined optimization principle can distinguish among them. Therefore the visible selector is the maximal canonical resolution available at the observable level. ∎
**Proposition 8.12 (Hidden entropy and fiber invariance)**
*Fix $m∈(0,1)$ and let $ν^⋆$ be the selected visible Bernoulli $(m)$ law from theorem 8.3. Then the hidden entropy rate is not constant on the observational fiber $E_Π(ν^⋆)$ . More precisely, for the stationary hidden laws $Q_λ,μ$ from corollary 8.6, one has
$$
H_hid(Q_λ,μ)=h_2(m)+(1-m)h_2(λ)+mh_2(μ),
$$
so distinct points of the same observational fiber can have different hidden entropy rates.*
* Proof*
By corollary 8.6, every law $Q_λ,μ$ belongs to the same observational fiber $E_Π(ν^⋆)$ . By proposition 8.9, the hidden entropy rate of the corresponding stationary hidden chain is
$$
H_hid(Q_λ,μ)=h_2(m)+(1-m)h_2(λ)+mh_2(μ).
$$
It therefore suffices to exhibit two parameter pairs in $(0,1)^2$ for which this value is different. Choose
$$
(λ,μ)=\Bigl(\frac{1}{2},\frac{1}{2}\Bigr) and (λ^\prime,μ^\prime)=\Bigl(\frac{1}{4},\frac{1}{4}\Bigr).
$$
Since the binary entropy function is strictly increasing on $(0,1/2)$ , one has
$$
h_2\Bigl(\frac{1}{4}\Bigr)<h_2\Bigl(\frac{1}{2}\Bigr). \tag{14}
$$
Hence
$$
\displaystyle H_hid(Q_1/4,1/4) \displaystyle=h_2(m)+(1-m)h_2\Bigl(\frac{1}{4}\Bigr)+mh_2\Bigl(\frac{1}{4}\Bigr) \displaystyle=h_2(m)+h_2\Bigl(\frac{1}{4}\Bigr) \displaystyle<h_2(m)+h_2\Bigl(\frac{1}{2}\Bigr) \displaystyle=H_hid(Q_1/2,1/2). \tag{14}
$$
Thus the hidden entropy rate is not constant on $E_Π(ν^⋆)$ . ∎
**Corollary 8.13 (Fiber invariance of visible selection)**
*Fix $m∈(0,1)$ and the selected visible law $ν^⋆$ from theorem 8.3. Then the canonical visible selector depends only on the visible feasible class and is therefore unchanged across the hidden observational fiber of $ν^⋆$ , whereas hidden-entropy maximization distinguishes between hidden laws inside that same fiber.*
* Proof*
The first statement is immediate from the definition of the selector: it is an optimizer of the visible entropy-rate functional on a visible feasible class, so once the visible feasible class is fixed, hidden realizations play no further role. In the present application, theorem 8.3 shows that the selected visible law is the Bernoulli $(m)$ law. By contrast, proposition 8.12 shows that the hidden entropy rate varies along the hidden observational fiber $E_Π(ν^⋆)$ . Therefore a hidden-entropy criterion distinguishes between hidden laws that are observationally indistinguishable at the selected visible level. This proves the stated contrast. ∎
**Proposition 8.14 (Fiber constancy of observable rules)**
*Fix a visible stationary law $ν$ and an observation map $Π$ . Let $S$ be any rule defined on hidden stationary laws with the property that whenever
$$
Π_\#Q=Π_\#Q^\prime,
$$
one has
$$
S(Q)=S(Q^\prime).
$$
Then $S$ is constant on the observational fiber $E_Π(ν)$ . In particular, any selection criterion determined only by the observation experiment and retained visible observables must be fiber-constant on each observational fiber.*
* Proof*
Let $Q,Q^\prime∈E_Π(ν)$ . By the definition of the observational fiber in section 2,
$$
Π_\#Q=ν=Π_\#Q^\prime.
$$
Therefore the defining property of $S$ gives
$$
S(Q)=S(Q^\prime).
$$
Since $Q,Q^\prime$ were arbitrary elements of $E_Π(ν)$ , the rule $S$ is constant on that fiber. For the final sentence, if a criterion is determined only by the observation experiment and retained visible observables, then any two hidden laws that induce the same retained visible specification are indistinguishable for that criterion. Hence the criterion satisfies the displayed implication above and is therefore fiber-constant. ∎
**Corollary 8.15 (Hidden entropy is not observable)**
*Fix $m∈(0,1)$ and the selected visible law $ν^⋆$ from theorem 8.3. Hidden-entropy maximization on the hidden observational fiber $E_Π(ν^⋆)$ is not an observable-determined criterion in the sense of proposition 8.14.*
* Proof*
By proposition 8.14, every observable-determined rule must be constant on each observational fiber. But proposition 8.12 shows that hidden entropy takes different values at different points of the single fiber $E_Π(ν^⋆)$ . Therefore hidden entropy cannot define an observable-determined criterion. ∎
**Proposition 8.16 (Hidden selectors cannot separate a fiber)**
*Fix a visible stationary law $ν$ and an observation map $Π$ . Assume that the observational fiber $E_Π(ν)$ contains at least two distinct hidden stationary laws. Let
$$
T:E_Π(ν)→P_stat(H^ℤ)
$$
be a rule such that
1. $T(Q)∈E_Π(ν)$ for every $Q∈E_Π(ν)$ ,
1. $T$ is observable-determined in the sense of proposition 8.14.
Then $T$ is constant on $E_Π(ν)$ . In particular, $T$ cannot separate distinct points of the fiber, equivalently, no observable-determined hidden selector can be injective on a non-singleton observational fiber.*
* Proof*
By proposition 8.14, every observable-determined rule is constant on each observational fiber. Therefore there exists some hidden stationary law $Q^†∈E_Π(ν)$ such that
$$
T(Q)=Q^† for every Q∈E_Π(ν).
$$
Since the fiber contains at least two distinct points, a constant map on that fiber cannot be injective. Hence $T$ cannot separate distinct points of $E_Π(ν)$ . ∎
**Corollary 8.17 (Application to the selected Bernoulli fiber)**
*Fix $m∈(0,1)$ and let $ν^⋆$ be the selected visible Bernoulli $(m)$ law from theorem 8.3. Then no observable-determined hidden selector can separate points of the observational fiber $E_Π(ν^⋆)$ . In particular, no such selector can recover a hidden law from the selected visible law in a one-to-one way.*
* Proof*
By corollary 8.6, the fiber $E_Π(ν^⋆)$ contains infinitely many distinct hidden stationary laws. The conclusion therefore follows immediately from proposition 8.16. ∎
**Proposition 8.18 (Factorization through visible laws)**
*Let
$$
S:P_stat(H^ℤ)→X
$$
be a map into an arbitrary set $X$ such that
$$
Π_\#Q=Π_\#Q^\prime \Longrightarrow S(Q)=S(Q^\prime).
$$
Then there exists a unique map
$$
\widetilde{S}:Π_\#\bigl(P_stat(H^ℤ)\bigr)→X
$$
such that
$$
S=\widetilde{S}∘Π_\#.
$$
In particular, every observable-determined rule factors through the visible stationary law.*
* Proof*
For each visible stationary law
$$
ν∈Π_\#\bigl(P_stat(H^ℤ)\bigr),
$$
choose any hidden stationary law $Q$ satisfying $Π_\#Q=ν$ , and define
$$
\widetilde{S}(ν):=S(Q).
$$
This is well-defined: if $Q^\prime$ is another hidden stationary law with $Π_\#Q^\prime=ν$ , then
$$
Π_\#Q=ν=Π_\#Q^\prime,
$$
so the hypothesis gives $S(Q)=S(Q^\prime)$ . Now let $Q∈P_stat(H^ℤ)$ . By construction,
$$
\widetilde{S}(Π_\#Q)=S(Q),
$$
so
$$
S=\widetilde{S}∘Π_\#.
$$
To prove uniqueness, suppose that another map $\widehat{S}$ satisfies
$$
S=\widehat{S}∘Π_\#.
$$
Let $ν$ lie in the image of $Π_\#$ , and choose $Q$ with $Π_\#Q=ν$ . Then
$$
\widetilde{S}(ν)=S(Q)=\widehat{S}(Π_\#Q)=\widehat{S}(ν).
$$
Hence $\widetilde{S}=\widehat{S}$ . ∎
**Remark 8.19 (Interpretive consequence)**
*The comparison above isolates the role of the selector. Hidden-identification problems ask whether the data determine a unique latent mechanism. Projection and lumpability problems ask when an aggregated visible process inherits a Markov structure from a hidden process. Externally parametrized maximum-entropy models begin by restricting attention to a prescribed functional family. The paper instead takes the observation experiment and retained visible observables as primitive, forms the corresponding visible feasible class, and then selects a canonical visible completion by entropy-rate maximization on that class.*
## 9 The Binary First-Order Case
I specialize theorems 4.1 and 4.4 to the binary first-order case.
Let $A=\{0,1\}$ and $r=1$ , and consider a stationary binary Markov chain with transition matrix
$$
P=\begin{pmatrix}1-a&a\\
b&1-b\end{pmatrix}, a,b∈[0,1].
$$
Its stationary mean $m=ℙ(X_t=1)$ satisfies
$$
m=(1-m)a+m(1-b).
$$
Equivalently,
$$
(1-m)a=m(1-b).
$$
Suppose the retained observable is the stationary mean $m$ . Then the feasible visible class consists of all stationary binary first-order laws with one-point marginal
$$
π=(1-m,m).
$$
By theorem 4.1, the unique entropy-rate selector on this feasible class is the i.i.d. Bernoulli $(m)$ law.
For explicit parametrization, introduce the flow coordinate
$$
q:=(1-m)a=m(1-b).
$$
Then
$$
a=\frac{q}{1-m}, 1-b=\frac{q}{m}, 0≤ q≤\min\{m,1-m\}.
$$
The entropy rate is therefore
$$
h(q)=(1-m) h_2\Bigl(\frac{q}{1-m}\Bigr)+m h_2\Bigl(\frac{q}{m}\Bigr),
$$
where
$$
h_2(p):=-p\log p-(1-p)\log(1-p)
$$
denotes binary entropy. Because $h_2$ is strictly concave on $(0,1)$ and the map $q↦(q/(1-m),q/m)$ is affine on the feasible interval, the function $q↦ h(q)$ is strictly concave on the feasible interval. Hence it has a unique maximizer. By corollary 4.2, that maximizer is attained at
$$
q^⋆=m(1-m),
$$
which yields
$$
a^⋆=m, b^⋆=1-m.
$$
Equivalently, the selected transition matrix is
$$
P^⋆=\begin{pmatrix}1-m&m\\
1-m&m\end{pmatrix}.
$$
The general gap functional from corollary 4.4 takes the form
$$
Δ=h_2(m)-\bigl[(1-m)h_2(a)+mh_2(1-b)\bigr].
$$
In the present binary first-order setting,
$$
Δ=I(X_0,X_1)≥ 0.
$$
Moreover,
$$
Δ=0 \Longleftrightarrow X_1⊥ X_0 \Longleftrightarrow a=m, b=1-m.
$$
Thus the gap vanishes exactly at the entropy maximizer, and it is strictly positive for every same-mean comparator with residual serial dependence.
A convenient persistence coordinate is
$$
ρ:=1-a-b.
$$
Under the fixed-mean constraint,
$$
a=m(1-ρ), b=(1-m)(1-ρ).
$$
Hence $ρ=0$ is exactly the canonical point selected by entropy rate, while nonzero $ρ$ measures residual serial structure inside the same retained-mean feasible class.
The binary case also illustrates the dynamical interpretation of the selector. The selected visible law is structureless at the retained observable level, yet it may still admit nontrivial hidden implementations that remain observationally invisible.
## 10 Conclusion
I have developed an entropy-rate selection framework on feasible classes determined by retained observables in the finite-state finite-memory setting. The paper gives global characterization results for the entropy maximizer, a local geometric theory on fixed-support faces, a hidden realization result that leaves the visible law unchanged, and a local empirical consistency theorem. The aliased hidden-state example shows that a maximizing visible completion need not resolve hidden underidentification.
## References
- [1] D. Blackwell, Comparison of Experiments, in Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, University of California Press, 1951, pp. 93–102.
- [2] J. C. C. McKinsey, A. W. Marshall, and M. K. Gardner, A simple proof of Blackwell’s ‘Comparison of Experiments’ theorem, Journal of Economic Theory 27 (1982), 439–443.
- [3] T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., John Wiley & Sons, 2006.
- [4] A. W. van der Vaart, Asymptotic Statistics, Cambridge University Press, 1998.
- [5] P. Billingsley, Probability and Measure, 3rd ed., John Wiley & Sons, 1995.