Data In Motion: 2020

Showing posts with label 2020. Show all posts

No More Spaghetti Flows

Spaghetti Flows

You may have heard of: https://en.wikipedia.org/wiki/Spaghetti_code. For Apache NiFi, I have seen some (and have done some of them in the past), I call them Spaghetti Flows.

Let's avoid them. When you are first building a flow it often meanders and has lots of extra steps and extra UpdateAttributes and random routes. This applies if you are running on-premise, in CDP or in other stateful NiFi clusters (or single nodes). The following video from Mark Payne is a must watch before you write any NiFi flows.

Apache NiFi Anti-Patterns with Mark Payne

https://www.youtube.com/watch?v=RjWstt7nRVY

https://www.youtube.com/watch?v=v1CoQk730qs

https://www.youtube.com/watch?v=JbUjYr6Kd3I

https://github.com/tspannhw/EverythingApacheNiFi

Do Not:

Do not Put 1,000 Flows on one workspace.
If your flow has hundreds of steps, this is a Flow Smell. Investigate why.
Do not Use ExecuteProcess, ExecuteScripts or a lot of Groovy scripts as a default, look for existing processors
Do not Use Random Custom Processors you find that have no documentation or are unknown.
Do not forget to upgrade, if you are running anything before Apache NiFi 1.10, upgrade now!
Do not run on default 512M RAM.
Do not run one node and think you have a highly available cluster.
Do not split a file with millions of records to individual records in one shot without checking available space/memory and back pressure.
Use Split processors only as an absolute last resort. Many processors are designed to work on FlowFiles that contain many records or many lines of text. Keeping the FlowFiles together instead of splitting them apart can often yield performance that is improved by 1-2 orders of magnitude.

Do:

Reduce, Reuse, Recycle. Use Parameters to reuse common modules.
Put flows, reusable chunks (write to Slack, Database, Kafka) into separate Process Groups.
Write custom processors if you need new or specialized features
Use Cloudera supported NiFi Processors
Use RecordProcessors everywhere
Read the Docs!
Use the NiFi Registry for version control.
Use NiFi CLI and DevOps for Migrations.
Run a CDP NiFi Datahub or CFM managed 3 or more node cluster.
Walk through your flow and make sure you understand every step and it’s easy to read and follow. Is every processor used? Are there dead ends?
Do run Zookeeper on different nodes from Apache NiFi.
For Cloud Hosted Apache NiFi - go with the "high cpu" instances, such as 8 cores, 7 GB ram.
same flow 'templatized' and deployed many many times with different params in the same instance
Use routing based on content and attributes to allow one flow to handle multiple nearly identical flows is better than deploying the same flow many times with tweaks to parameters in same cluster.
Use the correct driver for your database. There's usually a couple different JDBC drivers.
Make sure you match your Hive version to the NiFi processor for it. There are ones out there for Hive 1 and Hive 3! HiveStreaming needs Hive3 with ACID, ORC. https://community.cloudera.com/t5/Support-Questions/how-to-use-puthivestreaming/td-p/108430

Let's revisit some Best Practices:

https://medium.com/@abdelkrim.hadjidj/best-practices-for-using-apache-nifi-in-real-world-projects-3-takeaways-1fe6912101db

Get your Apache NiFi for Dummies. My own NiFi 101.

Here are a few things you should have read and tried before building your first Apache NiFi flow:

Also when in doubt, use Records! Use Record Processors and use pre-defined schemas, this will be easier to develop, cleaner and more performant. Easier, Faster, Better!!!

There are record processors for Logs (Grok), JSON, AVRO, XML, CSV, Parquet and more.

Look for a processor that has “Record” in the name like PutDatabaseRecord or QueryRecord.

Use the best DevOps processes, testing and tools.

Some newer features in 1.8, 1.9, 1.10, 1.11 that you need to use.

Advanced Articles:

Spaghetti is for eating, not for real-time data streams. Let's keep it that way.

If you are not sure what to do check out the Cloudera Community, NiFi Slack or the NiFi docs. Also I may have a helpful article here. Join me and my NiFi friends at virtual meetups for more in-depth NiFi, Flink, Kafka and more. We keep it interactive so you can feel free to ask questions.

Note: In this picture I am in Italy doing spaghetti research.

Commonly Used TCP/IP Ports in Streaming

Cloudera CDF and HDF Ports

NiFi and Friends
FLaNK Extended Stack

Note:

All of these ports can be changed by administrators or in version updates. Also if you are running Apache Knox like in Cloudera Data Platform Public Cloud, these ports may be changed or hidden. This is just based on a version of CDF I am running and defaults in. This does not include standard Cloudera ports for Cloudera Manager, Hadoop, Atlas, Ranger and other necessary and fun services.

Cloudera Flow Management (CFM Powered by Apache NiFi)

Cloudera NiFi HTTP: 8080 or 9090
Cloudera NiFi HTTPS: 8443 or 9443
Cloudera NiFi RIP Socket: 10443 or 50999
Cloudera NiFi Node Protocol: 11443
Cloudera NiFi Load Balancing: 6342
Cloudera NiFi Registry: 18080
Cloudera NiFi Registry SSL: 18433
Cloudera NiFi Certificate Authority: 10443

Cloudera Edge Flow Management (CEM Powered by Apache NiFi - MiNiFi)

Cloudera EFM HTTP: 10080
Cloudera EFM CoAP: 8989

Cloudera Stream Processing (CSP Powered by Apache Kafka)

Cloudera Kafka: 9092
Cloudera Kafka SSL: 9093
Cloudera Kafka Connect: 38083
Cloudera Kafka Connect SSL: 38085
Cloudera Kafka Jetty Metrics: 38084
Cloudera Kafka JMX: 9393
Cloudera Kafka MirrorMaker JMX: 9394
Cloudera Kafka HTTP Metric: 24042
Cloudera Schema Registry Registry: 7788
Cloudera Schema Registry Admin: 7789
Cloudera Schema Registry SSL: 7790
Cloudera Schema Registry Admin SSL: 7791
Cloudera Schema Registry Database (Postgresql): 5432
Cloudera SRM: 6669
Cloudera RPC: 8081
Cloudera SRM Rest: 6670
Cloudera SRM Rest SSL: 6671
Cloudera SMM Rest / UI: 9991
Cloudera SMM Manager: 8585
Cloudera SMM Manager SSL: 8587
Cloudera SMM Manager Admin: 8586
Cloudera SMM Manager Admin SSL: 8588
Cloudera SMM Service Monitor: 9997
Cloudera SMM Kafka Connect: 38083
Cloudera SMM Database (Postgresql): 5432

Cloudera Streaming Analytics (CSA Powered by Apache Flink)

Cloudera Flink Dashboard: 8082

References

Let's Query Kafka with Hive

Let's Query Kafka with Hive

I can hop into beeline and build an external Hive table to access my Cloudera CDF Kafka cluster whether it is in the public cloud in CDP DataHub, on-premise in HDF or CDF or in CDP-DC.

I just have to set my KafkaStorageHandler, Kafka Topic Name and my bootstrap servers (usually port 9092). Now I can use that table to do ELT/ELT for populating Hive tables or populating Kafka topics from Hive tables. This is a nice and easy way to do data engineering on the quick and easy.

This is a good item to augment CDP Data Engineering with Spark, CDP DataHub with NiFi, CDP DataHub with Kafka and KafkaStreams and various SQOOP or Python utilities you may have in your environment.

For real-time continuous queries on Kafka with SQL, you can use Flink SQL. https://www.datainmotion.dev/2020/05/flank-low-code-streaming-populating.html

Example Table Create

CREATE EXTERNAL TABLE <tableName>

(`uuid` STRING, `systemtime` STRING , `temperaturef` STRING , `pressure` DOUBLE,`humidity` DOUBLE, `lux` DOUBLE, `proximity` int, `oxidising` DOUBLE , `reducing` DOUBLE, `nh3` DOUBLE , `gasko` STRING,`current` INT, `voltage` INT ,`power` INT, `total` INT,`fanstatus` STRING)

STORED BY 'org.apache.hadoop.hive.kafka.KafkaStorageHandler'

TBLPROPERTIES

("kafka.topic" = "<TopicName>",

"kafka.bootstrap.servers"="<ServerName>:9092");

show tables;

describe extended kafka_table;

select *

from kafka_table;

I can browse my Kafka topics with Cloudera SMM to see what the data is and why I want to load or need to load.

For more information take a look at the documentation for Integrating Hive and Kafka at Cloudera below:

https://docs.cloudera.com/HDPDocuments/HDP3/HDP-3.1.5/integrating-hive/content/hive_kafka_query_table.html
https://docs.cloudera.com/runtime/7.0.3/integrating-hive-and-bi/topics/hive-kafka-integration.html
https://docs.cloudera.com/runtime/7.1.0/integrating-hive-and-bi/topics/hive-kafka-integration.html
https://docs.cloudera.com/runtime/7.1.0/integrating-hive-and-bi/topics/hive_ingest_kafka_data_into_hive.html

One Minute NiFi Tip: Calcite SQL Notes

NiFi Quick Tip on SQL

You sometimes have to cast, as fields aren't what you think they are. I have some temperatures that are stored as string, yeah I know let's yell at who did that. Maybe it was some lazy developer (Me?~??~?~?!!!). Let's just cast to a type that makes sense for math and comparisons. CAST is my friend.

SELECT *
FROM FLOWFILE
WHERE CAST(temperaturef as FLOAT) > 60

Apache NiFi (and lots of other awesome projects) use Apache Calcite for queries. So if you need some SQL help, always look here: https://calcite.apache.org/docs/reference.html

You can also include variables in your QueryRecord queries.

SELECT *
FROM FLOWFILE
WHERE CAST(temperaturef as FLOAT) >= (CAST(${predictedTemperature} as FLOAT) - 5)

There are wildcard characters that you may need to watch.

Underscore has special meaning. Also there often column names that are reserved words. I got a lot of columns coming from IoT often with names like timestamp, start, end and other ones used by SQL. Just put a `start` around it.

Watch those wildcards.

select * from flowfile where internal = false
                    and name not like '@_@_%' ESCAPE '@'

FLaNK: Low Code Streaming: Populating Kafka Topics with FlinkSQL Joins in Real-Time

FLaNK: Low Code Streaming: Populating Kafka Topics with FlinkSQL Joins in Real-Time

Then I can create my 3 tables. Two are the source ones to join and the third is the destination for my insert.

INSERT INTO global_sensor_events

SELECT

scada.uuid,

scada.systemtime ,

scada.temperaturef ,

scada.pressure ,

scada.humidity ,

scada.lux ,

scada.proximity ,

scada.oxidising ,

scada.reducing ,

scada.nh3 ,

scada.gasko,

energy.`current`,

energy.voltage ,

energy.`power` ,

energy.`total`,

energy.fanstatus

FROM energy,

scada

WHERE

scada.systemtime = energy.systemtime;

Examples

https://github.com/tspannhw/meetup-sensors/blob/master/flink-sql/

Assets / Scripts / DDL / SQL

https://github.com/tspannhw/FlinkSQLDemo

Flink Guide to SQL Joins

https://www.youtube.com/watch?v=5AuBlVRKQuo

Slides

https://www.slideshare.net/bunkertor/time-series-analysis-dataflow

Article on Joins

https://www.datainmotion.dev/2020/05/flink-sql-preview.html

Resources

Available Now: CFM 1.1.0

Cloudera Flow Management 1.1.0

Now with all the cool new features of Parameters, Predictive Monitoring, Rules Engine, Encrypted Repositories and SQL Reporting Tasks.

https://docs.cloudera.com/cfm/1.1.0/release-notes/topics/cfm-whats-new.html

Download it today: https://www.cloudera.com/downloads/cdf/cfm.html

Features are highlighted here: https://www.slideshare.net/bunkertor/introduction-to-apache-nifi-1114

https://www.datainmotion.dev/2020/03/using-nifi-cli-to-restore-nifi-flows.html

https://www.datainmotion.dev/2020/02/new-and-improved-its-nifi.html

Apache NiFi 1.11.4 Release Notes.

Using NiFi CLI to Restore NiFi Flows From Backups

Using NiFi CLI to Restore NiFi Flows From Backups

Please note, Apache NiFi 1.11.4 is now available for download.

https://cwiki.apache.org/confluence/display/NIFI/Release+Notes#ReleaseNotes-Version1.11.4

References:

#> registry list-buckets -u http://somesite.compute-1.amazonaws.com:18080

# Name Id Description
- ---- ------------------------------------ -----------
1 IOT 45834964-d022-4f4c-891f-695898e1e5f0 (empty)
2 IoT 250a5ae5-ced8-4f4e-8b3b-01eb9d47a0d9 (empty)
3 dev 46b7bab7-400f-44ae-a0e6-7340ff19c96f (empty)
4 iot c594d6bc-7413-4f6a-ba9a-50b8020eec37 (empty)
5 prod 0bf59d2e-1dd5-4d24-8aa0-0614bf991dc9 (empty)

#> registry create-flow -verbose -u http://somesite.compute-1.amazonaws.com:18080 -b 250a5ae5-ced8-4f4e-8b3b-01eb9d47a0d9 --flowName iotFlow

a5a4ac59-9aeb-416e-937f-e601ca8beba9

#> registry import-flow-version -verbose -u http://somesite.compute-1.amazonaws.com:18080 -f a5a4ac59-9aeb-416e-937f-e601ca8beba9 -i iot-1.json

#> registry list-flows -u http://ec2-35-171-154-174.compute-1.amazonaws.com:18080 -b 250a5ae5-ced8-4f4e-8b3b-01eb9d47a0d9

# Name Id Description
- ------- ------------------------------------ -----------
1 iotFlow a5a4ac59-9aeb-416e-937f-e601ca8beba9 (empty)

New and Improved: It's NiFi

Apache NiFi 1.11.3

http://nifi.apache.org/download.html

If you have downloaded anything after NiFi 1.10, please upgrade now. This has some major improvements and some fixes.

Release note highlights can be found here:
https://cwiki.apache.org/confluence/display/NIFI/Release+Notes#ReleaseNotes-Version1.11.3

I am running this now in Anaheim, and it's no Mickey Mouse upgrade. It's fast and nice.

Some of the more recent upgrades:

https://www.datainmotion.dev/2019/11/exploring-apache-nifi-110-parameters.html

For parameters and stateless, and ability to download a flow as JSON is worth the price of install.

See some more NiFi 1.11 features here: https://www.datainmotion.dev/2020/02/edgeai-google-coral-with-coral.html

EdgeAI: Google Coral with Coral Environmental Sensors and TPU With NiFi and MiNiFi (Updated EFM)

EdgeAI: Google Coral with Coral Environmental Sensors and TPU With NiFi and MiNiFi

Building MiNiFi IoT Apps with the new Cloudera EFM

It is very easy to build a drag and drop EdgeAI application with EFM and then push to all your MiNiFi agents.

Cloudera Edge Management CEM-1.1.1
Download the newest CEM today!

https://www.cloudera.com/downloads/cdf/cem.html

https://docs.cloudera.com/cem/1.1.1/release-notes/topics/cem-whats-new.html

NiFi Flow Receiving From MiNiFi Java Agent

In a cluster in my CDP-DC Cluster I consume Kafka messages sent from my remote NiFi gateway to publish alerts to Kafka and push records to Apache HBase and Apache Kudu. We filter our data with Streaming SQL.

We can use SQL to route, create aggregates like averages, chose a subset of fields and limit data returned. Using the power of Apache Calcite, Streaming SQL in NiFi is a game changer against Record Data Types including CSV, XML, Avro, Parquet, JSON and Grokable text. Read and write different formats and convert when your SQL is done. Or just to SELECT * FROM FLOWFILE to get everything.

We can see this flow from Atlas as we trace the data lineage and provenance from Kafka topic.

We can search Atlas for Kafka Topics.

From coral Kafka topic to NiFi to Kudu.

Details on Coral Kafka Topic

Examining the Hive Metastore Data on the Coral Kudu Table

NiFi Flow Details in Atlas

Details on Alerts Topic

Statistics from Atlas

See: https://www.datainmotion.dev/2020/02/connecting-apache-nifi-to-apache-atlas.html

Example Web Camera Image

Example JSON Record

[{"cputemp":59,"id":"20200221190718_2632409e-f635-48e7-9f32-aa1333f3b8f9","temperature":"39.44","memory":91.1,"score_1":"0.29","starttime":"02/21/2020 14:07:13","label_1":"hair spray","tempf":"102.34","diskusage":"50373.5 MB","message":"Success","ambient_light":"329.92","host":"coralenv","cpu":34.1,"macaddress":"b8:27:eb:99:64:6b","pressure":"102.76","score_2":"0.14","ip":"127.0.1.1","te":"5.10","systemtime":"02/21/2020 14:07:18","label_2":"syringe","humidity":"10.21"}]

Querying Kudu results in Hue

Pushing Alerts to Slack from NiFi

I am running on Apache NiFi 1.11.1 and wanted to point out a new feature. Download flow: Will download the highlighted flow/pgroup as JSON.

Looking at NiFi counters to monitor progress:

We can see how easy it is to ingest IoT sensor data and run AI algorithms on Coral TPUs.

Shell (coralrun.sh)

#!/bin/bash

DATE=$(date +"%Y-%m-%d_%H%M%S")

fswebcam -q -r 1280x720 /opt/demo/images/$DATE.jpg

python3 -W ignore /opt/demo/test.py --image /opt/demo/images/$DATE.jpg 2>/dev/null

Kudu Table DDL

https://github.com/tspannhw/table-ddl

Python 3 (test.py)

import time

import sys

import subprocess

import os

import base64

import uuid

import datetime

import traceback

import base64

import json

from time import gmtime, strftime

import math

import random, string

import time

import psutil

import uuid

from getmac import get_mac_address

from coral.enviro.board import EnviroBoard

from luma.core.render import canvas

from PIL import Image, ImageDraw, ImageFont

import os

import argparse

from edgetpu.classification.engine import ClassificationEngine

# Importing socket library

import socket

start = time.time()

starttf = datetime.datetime.now().strftime('%m/%d/%Y %H:%M:%S')

def ReadLabelFile(file_path):

with open(file_path, 'r') as f:

lines = f.readlines()

ret = {}

for line in lines:

pair = line.strip().split(maxsplit=1)

ret[int(pair[0])] = pair[1].strip()

return ret

# Google Example Code

def update_display(display, msg):

with canvas(display) as draw:

draw.text((0, 0), msg, fill='white')

def getCPUtemperature():

res = os.popen('vcgencmd measure_temp').readline()

return(res.replace("temp=","").replace("'C\n",""))

# Get MAC address of a local interfaces

def psutil_iface(iface):

# type: (str) -> Optional[str]

import psutil

nics = psutil.net_if_addrs()

if iface in nics:

nic = nics[iface]

for i in nic:

if i.family == psutil.AF_LINK:

return i.address

# /opt/demo/examples-camera/all_models

row = { }

try:

#i = 1

#while i == 1:

parser = argparse.ArgumentParser()

parser.add_argument('--image', help='File path of the image to be recognized.', required=True)

args = parser.parse_args()

# Prepare labels.

labels = ReadLabelFile('/opt/demo/examples-camera/all_models/imagenet_labels.txt')

# Initialize engine.

engine = ClassificationEngine('/opt/demo/examples-camera/all_models/inception_v4_299_quant_edgetpu.tflite')

# Run inference.

img = Image.open(args.image)

scores = {}

kCount = 1

# Iterate Inference Results

for result in engine.ClassifyWithImage(img, top_k=5):

scores['label_' + str(kCount)] = labels[result[0]]

scores['score_' + str(kCount)] = "{:.2f}".format(result[1])

kCount = kCount + 1

enviro = EnviroBoard()

host_name = socket.gethostname()

host_ip = socket.gethostbyname(host_name)

cpuTemp=int(float(getCPUtemperature()))

uuid2 = '{0}_{1}'.format(strftime("%Y%m%d%H%M%S",gmtime()),uuid.uuid4())

usage = psutil.disk_usage("/")

end = time.time()

row.update(scores)

row['host'] = os.uname()[1]

row['ip'] = host_ip

row['macaddress'] = psutil_iface('wlan0')

row['cputemp'] = round(cpuTemp,2)

row['te'] = "{0:.2f}".format((end-start))

row['starttime'] = starttf

row['systemtime'] = datetime.datetime.now().strftime('%m/%d/%Y %H:%M:%S')

row['cpu'] = psutil.cpu_percent(interval=1)

row['diskusage'] = "{:.1f} MB".format(float(usage.free) / 1024 / 1024)

row['memory'] = psutil.virtual_memory().percent

row['id'] = str(uuid2)

row['message'] = "Success"

row['temperature'] = '{0:.2f}'.format(enviro.temperature)

row['humidity'] = '{0:.2f}'.format(enviro.humidity)

row['tempf'] = '{0:.2f}'.format((enviro.temperature * 1.8) + 32)

row['ambient_light'] = '{0}'.format(enviro.ambient_light)

row['pressure'] = '{0:.2f}'.format(enviro.pressure)

msg = 'Temp: {0}'.format(row['temperature'])

msg += 'IP: {0}'.format(row['ip'])

update_display(enviro.display, msg)

# i = 2

except:

row['message'] = "Error"

print(json.dumps(row))

Source Code:

https://github.com/tspannhw/nifi-minifi-coral-env

Sensors / Devices / Hardware:

Humdity-HDC2010 humidity sensor
Light-OPT3002 ambient light sensor
Barometric-BMP280 barometric pressure sensor
PS3 Eye Camera and Microphone USB
Raspberry Pi 3B+
Google Coral Environmental Sensor Board
Google Coral USB Accelerator TPU