
Prompt
Traditionally, FL involves $L$ numbers of global training rounds, indexed by $\ell \in \mathcal{L}=\{1, 2, \cdots, L\}$, and within each $\ell$, each device trains local model with its local dataset. However, by considering devices' heterogeneities, such as communication, computation, and mobility, a DT--enabled model pruning FL framework is proposed to enhance the training performance. As shown in Fig.~1, in this paper, we group every $T_{\ell}$ numbers of global training rounds as one time slot, denoted as $t$. Suppose that there exist $T$ numbers of time slots, and within each time slot $t \in \mathcal{T}=\{1, 2, \cdots, T\}$, we perform pruning global model to preserve the structured sparsity and model performance. In a global training round $\ell$, we further prune parameters of the pruned global model.In our framework, at the beginning of each federated learning round $\ell$, the edge server makes initial pruning rate and bandwidth allocation decisions for each device based on the computing frequency and location reported by the device, and sends these decisions back to the devices. Devices periodically report their current computing frequency and location to the digital twin layer and the edge server. Upon receiving these updates, the digital twin layer immediately refreshes the virtual replicas and adjusts the edge server's pruning rate and bandwidth allocation decisions, thereby achieving rapid local policy updates. Simultaneously, the digital twin layer validates and optimizes these decisions based on the constructed global virtual environment. Ultimately, the optimized global decisions are returned to the devices within the same global federated learning round}. The procedures of proposed DT--enabled model pruning FL can be described as follows. 1) \textbf{Determinations of Model Pruning Ratios}: At the beginning of time slot $\ell$, each device uploads its up--to--date states, such as computing capacity, to its DT. The DT will determine $\Psi_{k}^{t,\ell}$ and $\Phi_{k}^{t,\ell}$, which will be detailed later. Then, based on $\Psi_{k}^{t,\ell}$, the DT will prune the global model $\textbf{w}$ to obtain ${\textbf{w}}_{k}^{t,\ell}$, and train ${\textbf{w}}_{k}^{t,\ell}$ on its synthetic dataset $\hat{\mathcal{D}}_k$.2) \textbf{Local Training}: At time slot $\ell$, device $k$ trains the received model ${\textbf{w}}_{k}^{t,\ell}$ on $\mathcal{D}_k$ for $\varsigma_k$ numbers of local iterations to obtain $\tilde{{\textbf{w}}}_{k}^{t,\ell}$.3) \textbf{Locally Model Pruning}: Based on received $\Phi_{k}^{t,\ell}$, device $k$ prunes the model $\tilde{{\textbf{w}}}_{k}^{t,\ell}$.4) \textbf{Model Upload}: Once completion of local model training, devices should transmit their models to DTs.5) \textbf{Global Aggregation}: Upon the receipt of all local models, we implement the following aggregation scheme at each slot $\ell$.6) \textbf{Model Distribution}: At the next global training round, the edge server transfers the aggregated model to DTs, and a next learning process is launched. 注意:将edge server和DT画在一起,device不要与edge server和DT画在一起
GPT Image-2 - Text to Image
Settings
Create Something Similar
The generator below preselects the model used for this record so you can remix the idea faster.