0
点赞
收藏
分享

微信扫一扫

Flink 流处理 API _Source

1 从集合读取数据 
// 定义样例类,传感器id,时间戳,温度
case class SensorReading(id: String, timestamp: Long, temperature: Double)

object Sensor {
def main(args: Array[String]): Unit = {
val env = StreamExecutionEnvironment.getExecutionEnvironment
val stream1 = env
.fromCollection(List(
SensorReading("sensor_1", 1547718199, 35.80018327300259),
SensorReading("sensor_6", 1547718201, 15.402984393403084),
SensorReading("sensor_7", 1547718202, 6.720945201171228),
SensorReading("sensor_10", 1547718205, 38.101067604893444)
))

stream1.print("stream1:").setParallelism(1)
env.execute()
}
}

2 从文件读取数据

val stream2 = env.readTextFile("YOUR_FILE_PATH")


3 以 kafka 消息队列的数据作为来源
需要引入 kafka 连接器的依赖: pom.xml

<!--
​​https://mvnrepository.com/artifact/org.apache.flink/flink-connector-kafka-0.11 ​​
-->
<dependency>
groupId>org.apache.flink</groupId>
artifactId>flink-connector-kafka-0.11_2.11</artifactId>
version>1.7.2</version>
</dependency>


具体代码如下:
val properties = new Properties()
properties.setProperty("bootstrap.servers", "localhost:9092") properties.setProperty("group.id", "consumer-group") properties.setProperty("key.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer") properties.setProperty("value.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer")
properties.setProperty("auto.offset.reset", "latest")
val stream3 = env.addSource(new FlinkKafkaConsumer011[String]("sensor", new
SimpleStringSchema(), properties))
4 自定义 Source
除了以上的 source 数据来源,我们还可以自定义 source。需要做的,只是传入
一个 SourceFunction 就可以。具体调用如下:

val stream4 = env.addSource( new MySensorSource() )

我们希望可以随机生成传感器数据,MySensorSource 具体的代码实现如下:
class MySensorSource extends SourceFunction[SensorReading]{

// flag: 表示数据源是否还在正常运行
var running: Boolean = true

override def cancel(): Unit = { running = false
} override def run(ctx: SourceFunction.SourceContext[SensorReading]): Unit
= {
// 初始化一个随机数发生器
val rand = new Random()
var curTemp = 1.to(10).map( i => ( "sensor_" + i, 65 + rand.nextGaussian() * 20 ) )

while(running){
// 更新温度值
curTemp = curTemp.map( t => (t._1, t._2 + rand.nextGaussian() )
)
// 获取当前时间戳
val curTime = System.currentTimeMillis()

curTemp.foreach(
t => ctx.collect(SensorReading(t._1, curTime, t._2))
)
Thread.sleep(100)
}
}
}

举报

相关推荐

0 条评论