Efficient Multiuser AI Downloading via Reusable Knowledge Broadcasting

Wu, Hai; Zeng, Qunsong; Huang, Kaibin

File Download

There are no files associated with this item.

Links for fulltext

(May Require Subscription)

Publisher Website: 10.1109/TWC.2024.3373015
Scopus: eid_2-s2.0-85187981664
Find via

Supplementary

Citations:
- Scopus: 0
Appears in Collections:
- Electrical & Electronic Engineering: Journal/Magazine Articles

Article: Efficient Multiuser AI Downloading via Reusable Knowledge Broadcasting

Title	Efficient Multiuser AI Downloading via Reusable Knowledge Broadcasting
Authors	Wu, Hai Zeng, Qunsong Huang, Kaibin
Keywords	Artificial intelligence broadcasting Broadcasting Computational modeling Edge AI in-situ model downloading knowledge reuse Measurement power control Protocols Servers Task analysis
Issue Date	12-Mar-2024
Publisher	Institute of Electrical and Electronics Engineers
Citation	IEEE Transactions on Wireless Communications, 2024 How to Cite? DOI: http://dx.doi.org/10.1109/TWC.2024.3373015
Abstract	For the sixth-generation (6G) mobile networks, in-situ model downloading has emerged as an important use case to enable real-time adaptive artificial intelligence (AI) on edge devices. However, the simultaneous downloading of diverse and high-dimensional models to multiple devices over wireless links presents a significant communication bottleneck. To overcome the bottleneck, we propose the framework of model broadcasting and assembling (MBA), which represents the first attempt on leveraging reusable knowledge , referring to shared parameters among tasks/models, to enable parameter broadcasting to reduce communication overhead or latency. The MBA framework comprises two key components. The first, the MBA protocol, defines the system operations including parameter selection from an AI library, power control for broadcasting, and model assembling at devices. The protocol features the use of Shapley value as a metric for measuring parameters’ reusability. The second component is the joint design of parameter-selection-and-power-control (PS-PC), which provides guarantees on devices’ model performance and aims to minimize the downloading latency. The corresponding optimization problem is simplified by decomposition into the sequential PS and PC sub-problems without compromising its optimality. The PS sub-problem is solved efficiently by designing two efficient algorithms. On one hand, the low-complexity algorithm of greedy parameter selection features the construction of task-oriented candidate model sets and a greedy selection metric for choosing the sets of model blocks for broadcasting, both of which are designed under the criterion of maximum reusable knowledge among tasks. On the other hand, the optimal tree-search algorithm gains its efficiency via the proposed construction of a compact binary tree pruned using model architecture constraints and an intelligent branch-and-bound search on the tree that fathoms nodes via solving a linear program that integer-relaxes the PS sub-problem. Last, given optimal PS, the optimal PC policy is derived in closed form by transforming the PC sub-problem into the conventional problem of energy-efficient transmission. Through extensive experiments conducted on real-world datasets, our results demonstrate the substantial reduction in downloading latency achieved by the proposed MBA design compared to traditional unicasting-based model downloading.
Persistent Identifier	http://hdl.handle.net/10722/344698
ISSN	1536-1276 2023 Impact Factor: 8.9 2023 SCImago Journal Rankings: 5.371

DC Field	Value	Language
dc.contributor.author	Wu, Hai	-
dc.contributor.author	Zeng, Qunsong	-
dc.contributor.author	Huang, Kaibin	-
dc.date.accessioned	2024-08-02T04:43:46Z	-
dc.date.available	2024-08-02T04:43:46Z	-
dc.date.issued	2024-03-12	-
dc.identifier.citation	IEEE Transactions on Wireless Communications, 2024	-
dc.identifier.issn	1536-1276	-
dc.identifier.uri	http://hdl.handle.net/10722/344698	-
dc.description.abstract	<p>For the sixth-generation (6G) mobile networks, in-situ model downloading has emerged as an important use case to enable real-time adaptive artificial intelligence (AI) on edge devices. However, the simultaneous downloading of diverse and high-dimensional models to multiple devices over wireless links presents a significant communication bottleneck. To overcome the bottleneck, we propose the framework of model broadcasting and assembling (MBA), which represents the first attempt on leveraging reusable knowledge , referring to shared parameters among tasks/models, to enable parameter broadcasting to reduce communication overhead or latency. The MBA framework comprises two key components. The first, the MBA protocol, defines the system operations including parameter selection from an AI library, power control for broadcasting, and model assembling at devices. The protocol features the use of Shapley value as a metric for measuring parameters’ reusability. The second component is the joint design of parameter-selection-and-power-control (PS-PC), which provides guarantees on devices’ model performance and aims to minimize the downloading latency. The corresponding optimization problem is simplified by decomposition into the sequential PS and PC sub-problems without compromising its optimality. The PS sub-problem is solved efficiently by designing two efficient algorithms. On one hand, the low-complexity algorithm of greedy parameter selection features the construction of task-oriented candidate model sets and a greedy selection metric for choosing the sets of model blocks for broadcasting, both of which are designed under the criterion of maximum reusable knowledge among tasks. On the other hand, the optimal tree-search algorithm gains its efficiency via the proposed construction of a compact binary tree pruned using model architecture constraints and an intelligent branch-and-bound search on the tree that fathoms nodes via solving a linear program that integer-relaxes the PS sub-problem. Last, given optimal PS, the optimal PC policy is derived in closed form by transforming the PC sub-problem into the conventional problem of energy-efficient transmission. Through extensive experiments conducted on real-world datasets, our results demonstrate the substantial reduction in downloading latency achieved by the proposed MBA design compared to traditional unicasting-based model downloading.<br></p>	-
dc.language	eng	-
dc.publisher	Institute of Electrical and Electronics Engineers	-
dc.relation.ispartof	IEEE Transactions on Wireless Communications	-
dc.subject	Artificial intelligence	-
dc.subject	broadcasting	-
dc.subject	Broadcasting	-
dc.subject	Computational modeling	-
dc.subject	Edge AI	-
dc.subject	in-situ model downloading	-
dc.subject	knowledge reuse	-
dc.subject	Measurement	-
dc.subject	power control	-
dc.subject	Protocols	-
dc.subject	Servers	-
dc.subject	Task analysis	-
dc.title	Efficient Multiuser AI Downloading via Reusable Knowledge Broadcasting	-
dc.type	Article	-
dc.identifier.doi	10.1109/TWC.2024.3373015	-
dc.identifier.scopus	eid_2-s2.0-85187981664	-
dc.identifier.eissn	1558-2248	-
dc.identifier.issnl	1536-1276	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Article: Efficient Multiuser AI Downloading via Reusable Knowledge Broadcasting

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats